# teerth.blog: full text > Long-form essays on the projects, with the maths, the proofs and the numbers. By Teerth Sharma (https://teerth.dev). > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # resolvent > A JEPA planner ranks candidate futures one at a time. When they share one error, the right pick depends on the whole candidate set. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/resolvent - Repository: https://github.com/teerthsharma/resolvent - Project site: https://teerth.dev/resolvent/ - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "resolvent", 2026, https://teerth.blog/resolvent - Languages: Python, Lean 4 - Question: When every candidate's prediction shares one error, how much can a set operator over the candidates recover, and where does it stop helping? - Headline result: 0.732 / 0.703 / 0.787 normalised score, D-JEPA-spec operator at ε = 4, K = 63, full shared error, 3 seeds (control: D-JEPA as specified (ε = 0.2): 0.352 / 0.354 / 0.358) ## What it is A latent world-model planner decides by ranking. From one start it proposes K candidate action sequences, rolls each one forward in latent space, and executes the candidate whose predicted future lies nearest the goal embedding. The rule looks at one candidate at a time: candidate $k$ is scored by $\lVert \hat z_k - z_g \rVert$ and by nothing else. That rule is the right one when the error separating prediction from outcome is independent for each candidate. In that case the Bayes pick and the distance pick coincide; the repository's test `B0` checks that they agree on at least 97% of starts when no error is shared. But a large part of a planner's error is not per-candidate. Every candidate of a start is rolled out from the same misestimated start, so that error moves all K predicted futures together. Once the error is a common, unknown translation of the whole set, the question "which candidate is best?" stops being a question about each candidate's distance and becomes a question about the shape of the set. resolvent is my attempt to measure that question exactly. It builds small beds where the truth is known, where the share of error that is common to all candidates can be dialled from none to all of it, and where the Bayes-optimal pick can be computed by Monte Carlo. Between the latent-distance floor and that Bayes ceiling it places a family of rankers: a pointwise head that sees each candidate alone, a one-hop set head, a resolvent set head that solves $(I - A)^{-1}$ over the candidates, and a reimplementation of D-JEPA's bounded relational operator. The README states the scope in a sentence I still hold to: "This repository measures where that helps D-JEPA's decision-local ranking, where it does not, and what it costs." It builds on D-JEPA and makes no novelty claim. The resolvent read itself is older than the planning question. It comes from an earlier causal attention programme in the same repository, a family in which softmax attention, unnormalised-kernel attention and the exact path product of a Markov chain are settings of one head. That programme's Lean proofs live in the same tree, and this essay uses the parts of them that bear on the set operator. The smallest bed, `shift`, makes the problem concrete with K = 4 candidates in D = 2 dimensions. The planner observes a start estimate $y$. The true start is $$ s = y - \sigma \xi, \qquad \xi \sim \mathcal N(0, I_D), \qquad\text{so}\quad s \mid y \sim \mathcal N(y, \sigma^2 I_D). $$ Candidates aim at the goal $g = 0$ from the estimate, $a_k = -y + \rho\, r_k$ with $r_k \sim \mathcal N(0, I_D)$, and execute to $$ z_k = s + a_k + \sigma_e \eta_k = \rho\, r_k - \sigma \xi + \sigma_e \eta_k, \qquad k^\star = \arg\min_k \lVert z_k \rVert. $$ The term $\sigma\xi$ is the same for every $k$. The plug-in rule and the Bayes rule are $$ \hat k_{\text{dist}} = \arg\min_k \lVert \hat z_k \rVert, \qquad \hat k_{\text{Bayes}} = \arg\max_k\ \Pr\!\left(k = k^\star \mid y, a_{1:K}\right), $$ with the probability estimated from $M = 1024$ posterior draws of $(\xi, \eta)$ that are common to all K candidates. **Figure 1.** Four candidates around the goal, with a shared start error of scale σ that translates all of them together. Each candidate's cell is shaded by the probability that it is the true best, estimated live from common posterior draws (the Bayes rule of the third equation). The grey ring marks the latent-distance pick, the green ring the Bayes pick. As σ grows the distance pick becomes the candidate inside the hull, whose cell the shared error almost never lands in. A mix slider moves error from shared to per-candidate; at full per-candidate error the two picks agree again, and the cell shading is hidden because the cells are no longer the answer. A run button replays N = 1500 starts and reports the hit rate of both picks; ticks on the probability bars mark the σ → ∞ hull shares. The lower panel plots the repository's recorded gap between Bayes and distance against σ beside the live recomputation. Colour key: baseline: latent-distance pick (the plug-in rule); proof: Bayes pick, the exact-truth oracle that knows σ; ticks mark the σ → ∞ hull share; measured: gap recorded in the repository (README.md:315), filled dots; hollow rings are the live recomputation; structure: candidates, rings and draws computed live; seq: probability that a candidate is the best one. As the shared error grows past the spread of the candidates, the chance that candidate $k$ is best tends to the exterior angle of its Voronoi cell. A candidate inside the convex hull of the set has exterior angle zero, and that interior candidate is exactly the one nearest-to-goal favours when the planner aims every candidate at the goal. Distance then drops below chance. A head that scores each candidate from its own prediction cannot represent a rule that depends on the others; a head that passes messages across the set can. That is the whole motivation for putting a set operator, and in particular a resolvent, between the predictor and the decision. ## What it can do The first thing the repository does is measure the size of the problem. On bed `shift` (K = 4, D = 2, $\rho = 1$, $\sigma_e = 0.02$, three seeds of 4,000 held-out starts each), the gap in top-1 hit rate between the Bayes pick and the latent-distance pick is 0.000, 0.024, 0.149, 0.213 and 0.241 at $\sigma$ = 0.1, 1, 3, 10 and 30. These five numbers are pinned to three decimals by `tests/resolvent/test_reproduce.py`; I did not re-run them for this essay. **Measured: 0.213** top-1 hit gap, Bayes pick minus latent-distance pick, bed shift, σ = 10, K = 4. Control: latent-distance pick on the same starts; the same gap is 0.000 at σ = 0.1. n = 3 seeds × 4,000 held-out starts; Bayes by M = 1,024 common draws. Source: README.md:315 @ 3b1b7c1, pinned by tests/resolvent/test_reproduce.py; recorded run, not re-run here. The second thing is to ask what closes that gap. At $\sigma = 1$, with a learned JEPA MLP predictor and 36 training epochs, the resolvent set head closes 0.947 of the gap, the one-hop set head 0.925, and the pointwise head, given the same inputs but no set interaction, 0.028. Closure here is (head − distance) / (Bayes − distance), and the denominator at $\sigma = 1$ is the 0.024 gap above, so per-seed closures are noisy: the resolvent's three seeds are 0.990, 0.859 and 0.992. The comparison that matters is not resolvent against one-hop. Their hit rates differ by 0.0002, in the one-hop head's favour. The comparison that matters is any set head against the pointwise control. **Measured: 0.947** share of the Bayes-minus-distance gap closed by the resolvent set head, σ = 1, 36 epochs. Control: pointwise head with the same inputs and no set interaction: 0.028. n = 3 seeds (per seed 0.990 / 0.859 / 0.992). Source: README.md:317-319 @ 3b1b7c1; experiments/shift/results_cell_sigma1_ep36_point-hop1-resolvent.json. The larger bed, `dial`, has K = 63 candidates in D = 8, a known five-step map $z' = 1.1\,Qz + 0.5\tanh(Cz + Ba)$, a learned MLP predictor, three seeds of 20,000 held-out starts, and a Bayes pick estimated from M = 2,048 draws under the true dynamics. Its dial is the fraction $f$ of start-error variance that is drawn once and shared by all candidates. Because the Bayes hit at K = 63 is small (about 0.083 to 0.086 against a blind rate of $1/K \approx 0.0159$), the repository reports a normalised score: $$ \mathrm{NS} = \frac{\mathrm{hit} - 1/K}{\mathrm{hit}_{\text{Bayes}} - 1/K}. $$ **Measured: 0.732 / 0.703 / 0.787** NS of the D-JEPA-spec operator with ε = 4 and LIN1 base ranks (dj4L), bed dial, f = 1, σ/ρ = 3. Control: D-JEPA as specified (ε = 0.2): 0.352 / 0.354 / 0.358; latent-distance floor 0.338 / 0.343 / 0.335. n = 3 seeds × 20,000 held-out starts; every learned arm sees the same 12-d token and a 4,000-step budget at 67,805–69,377 parameters. Source: README.md:328-337 @ 3b1b7c1; experiments/dial/r2/results/table.json; recorded run, not re-run here. On the same cell, the one-hop set head scores 0.501, 0.467 and 0.497. NS of 0.73 is a normalised score and not a 73% hit rate; the hit rate of the best learned arm is 0.065, 0.065 and 0.070, derived from the NS and the Bayes hit of 0.0833, 0.0856 and 0.085. **Figure 2.** The repository's recorded results in its own units. Left: bed shift at σ = 1, K = 4, hit rate per seed for the Bayes pick, the resolvent and one-hop set heads, the pointwise control and latent distance. Right: bed dial at K = 63 with all error shared, NS per seed for Bayes restricted to D-JEPA's reach, the ε = 4 operator, the one-hop head, D-JEPA as specified, the pointwise control and the distance floor. A units control switches between hit rate, NS and closure. An ε control draws the reach window of D-JEPA's bound on the normalised-rank axis and says when the bound constrains nothing. Bars are means, dots are seeds; nothing in this figure is trained in the browser. Colour key: measured: learned set arms, copied from the repository result files; baseline: latent distance and D-JEPA as specified (ε = 0.2); proof: Bayes ceiling and Bayes restricted to the 2ε reach (oracle references); the shaded band is the closable room from distance to Bayes, and the window under panel A is the reach; structure: controls: pointwise head and label-shuffled null. Three other results belong in this section. On the third bed, `torus`, the planner has to land in a goal ball after T steps of the Chirikov standard map on the 2-torus, with K = 8. A call-accounted race (B2, a Bernstein stopping rule) reaches the full rollout's quality within 0.01 NS on all 15 seed-lead rows while spending fewer predictor calls per decision; the saving is 4.52× to 7.89×. A sharper version, B4 at a budget of 656 calls and T = 32, scores NS′ of 0.99101, 0.99213 and 0.99021 on fresh seeds 14, 15 and 16 against the 128-member full rollout. **Measured: 4.52× – 7.89×** predictor calls saved per decision by the B2 race, bed torus, T ∈ 12, 16, 20, 24, 32. Control: full rollout, K(M + 1)T calls; the race stays within 0.01 NS of it on all 15 seed-lead rows. n = 15 seed-lead rows (seeds 3–5 × 5 leads); tuned on seeds 9 and 10. Source: README.md:360 @ 3b1b7c1; experiments/torus/r2/BAR.md:104. Finally, the repository contains a verifier, `daedalus/`, that grades candidate ranking code. Candidate code runs in a separate sandbox process that sees no evaluation labels and may not read outside its own directory. In round 2 it rejected 22 of 22 planted cheats with 0 errors, 21 of them at their registered stage, while an honest control passed. In round 4 it rejected 7 of 7 planted ranker cheats, 6 of them at the named stage, with both controls admissible. **Measured: 22 / 22** planted cheats rejected by the daedalus verifier, round 2. Control: honest control passes; 0 errors. n = 22 planted cheats; 21 of 22 rejected at their registered stage. Source: daedalus/results/r2/m0_r2c_redteam.json @ 3b1b7c1; machine WIN-16QAL06O9GB, Python 3.11.9. No upstream contribution is linked to this project in the lineage the blog records, so there is no upstream card here. ## How it was made The construction has three layers: the exact-truth beds that make the Bayes rule computable, the set operators that are trained against it, and the proofs underneath the operators. > **Problem: The shared-error ranking problem.** > > Given predictions $\hat z_1, \dots, \hat z_K$ of one start's candidates, and an outcome error that is partly a single draw common to all K, choose the candidate most likely to be the true best. The plug-in rule $\hat k_{\text{dist}}$ ignores the common part. The Bayes rule $\hat k_{\text{Bayes}}$ integrates over it. ### The hull limit Expand $\lVert \rho r_k - \sigma\xi \rVert^2$ and the cross term $-2\rho\sigma\, r_k^\top \xi$ dominates for large $\sigma$, so the true best tends to $\arg\max_k \xi^\top r_k$. For isotropic $\xi$ the direction is uniform on the circle, which gives $$ \Pr(k = k^\star) \;\to\; \frac{1}{2\pi}\left|\{u \in S^1 : k = \arg\max_j u^\top r_j\}\right|, $$ the share of directions in which candidate $k$ is the furthest point of the set. The repository computes this share in `hull_angle_share` (`resolvent/shared_error.py:64-74`) with 4,096 directions. It is zero for any candidate strictly inside the convex hull. Under a large shared translation, the Bayes rule is a convex-hull property. **Figure 3.** One candidate set at six values of the shared error σ, from 0.3 to the limit. In each panel the Voronoi cells are shaded by the probability that the candidate is best, computed by polar quadrature with no Monte Carlo. As σ grows, the mass of the interior candidate drains to zero and the three hull vertices take shares equal to their exterior angles, drawn as wedges in the last panel. The grey ring marks the distance pick, which is the interior point. A curve below tracks the largest difference between the live masses and the limiting shares, with the interior candidate's mass beside it; a scrubber sets σ, candidates can be dragged and the set can be swapped for a square or a random draw. Colour key: seq: probability that the candidate is best at this σ; proof: the σ → ∞ limit: exterior-angle share of each hull vertex; baseline: latent-distance pick; structure: candidates, and the live mass of the interior candidate; measured: hit rates recorded at σ = 10. ### D-JEPA's bounded correction and its reach D-JEPA (Liu et al., arXiv 2609.24749) closes part of what it calls the decision-local gap with a bounded, permutation-equivariant relational operator over the candidates. I reimplemented it from the paper's equations, not from the released code, in `resolvent/djepa.py`: $$ h_i = \mathrm{TF}(\mathrm{enc}(v_i)), \qquad \delta_i = \varepsilon \tanh\!\big(W_{\text{up}} \tanh(W_{\text{down}} h_i)\big), \qquad s_i = b_i + \delta_i, $$ where $b_i$ is the normalised base rank and lower scores are better. Since $|\delta_i| \le \varepsilon$, the selected candidate $\hat i = \arg\min_i s_i$ satisfies the paper's Prop 2 / Cor 1: $$ b_{\hat i} \;\le\; \min_i b_i + 2\varepsilon . $$ The forward pass is short: ```python def forward(self, v, base, pad=None): """v: (B, K, nin) tokens; base: (B, K); pad: (B, K) bool, True = padding. Returns (s, delta).""" h = self.tf(self.enc(v), src_key_padding_mask=pad) delta = self.eps * torch.tanh(self.up(torch.tanh(self.down(h)))).squeeze(-1) return base + delta, delta ``` `resolvent/djepa.py:33-37 @ 3b1b7c1`. The up projection is zero-initialised (`djepa.py:29-30`), so an untrained head returns the base ranking exactly, and every departure from latent distance is learned. The bound in the second equation also gives a learning-free ceiling: the Bayes rule restricted to candidates within $2\varepsilon$ of the base minimum is the best any operator obeying it can do. When $2\varepsilon$ exceeds the base range, the bound constrains nothing. **Figure 4.** D-JEPA's reach. The reader sets ε, K, the gain a (t = tanh a) and the head scale, and feeds the final tanh random or adversarial values; the selected candidate's base rank never exceeds the base minimum plus 2ε. The green line is the ε a target rank needs in order to win; ranks with base rank below 2εt form the reach window. When 2εt reaches the full rank range the window fills the axis and the bound constrains nothing. A K = 4 panel runs the Bayes pick live against the pick restricted to the reach. The transformer that produces h is not simulated. Colour key: proof: the 2ε bound and the ε each rank needs to win; baseline: D-JEPA as specified, ε = 0.2; structure: live draws of the correction, the ε·t level and the K = 4 Bayes-rank bars; measured: repository numbers on the same quantities; withdrawn: ε = 4, where the bound is vacuous. On bed `shift` at $\sigma = 10$, the best hit reachable under the bound at $\varepsilon = 0.2$ is 0.289, against a Bayes hit of 0.381. On bed `dial`, Bayes restricted to base ranks within $2\varepsilon = 0.4$ scores 0.944, 0.944 and 0.970 NS, while the learned $\varepsilon = 0.2$ operator scores 0.352 to 0.358. Its tanh correction runs at a mean $|\delta|/\varepsilon$ of 0.792. The README puts it in one line: "Reach is not the wall; saturation is." Widening $\varepsilon$ to 0.5, 1 or 4 with everything else fixed lifts NS to the same level. ### The resolvent on a candidate set A causal resolvent, the kind used over a sequence, is nilpotent off the diagonal and cannot reach a pole. On an unordered set the coupling has to be diagonal-free and must not depend on the listing order. Such a coupling is no longer nilpotent, and the bound on it becomes load-bearing. > **Definition: Signed candidate-set resolvent.** > > For queries $q$, keys $k$ and values $v$ over K candidates, and $\rho < 1$, > > $$ > A = \rho\, \frac{\operatorname{offdiag}(q k^\top)}{\operatorname{rowL1} + \epsilon}, \quad > \lVert A \rVert_\infty \le \rho < 1, \quad > \text{out} = (I - A)^{-1} v = \sum_{h \ge 0} A^h v, \quad > \lVert (I - A)^{-1} \rVert_\infty \le \frac{1}{1 - \rho}. > $$ > > $(A^h)_{ij}$ sums the weights of every length-$h$ message path from candidate $j$ to candidate $i$, so one solve reads the path structure of the whole set. $\rho \ge 1$ is refused. The code is the definition, with two engineering decisions visible in it: ```python def build_A(q, k, rho=0.9, mask=None, eps=1e-3): """Signed, diagonal-free, row-L1-normalised coupling. eps floors the row L1: without it |dA/dq| ~ 1/|q| grows without bound as the logit scale goes to zero.""" if not rho < 1.0: raise ValueError(f"rho={rho!r}: off the candidate axis A is not nilpotent; rho >= 1 admits a pole") ct = torch.promote_types(q.dtype, torch.float32) q, k = q.to(ct), k.to(ct) K = q.shape[-2] w = q @ k.transpose(-1, -2) keep = ~torch.eye(K, dtype=torch.bool, device=q.device) if mask is not None: keep = keep & mask[..., :, None] & mask[..., None, :] w = w.masked_fill(~keep, 0.0) return rho * w / (w.abs().sum(-1, keepdim=True) + eps).clamp_min(torch.finfo(ct).tiny) ``` `resolvent/resolvent.py:19-32 @ 3b1b7c1`. The first decision is precision: the coupling is formed in at least float32, because there is no fp16 or bf16 LU and a large $q \cdot k$ overflows fp16 to infinity, after which the row normalisation is NaN. "Low precision is promoted, not trusted." The second is the solve. `set_resolvent` builds $I - A$ and calls a dense `torch.linalg.solve`, because a triangular solve applied to a non-causal matrix silently reads only its lower triangle. Masked candidates are removed on both axes and output 0, and an all-masked row returns 0 instead of NaN. A truncated alternative, `neumann_resolvent`, sums $v + Av + \dots + A^H v$; from $\lVert A \rVert_\infty \le \rho$ and the geometric tail, its error is at most $$ \frac{\rho^{H+1}}{1-\rho}\,\max|v| . $$ **Figure 5.** The signed set resolvent run live on the reader's candidates. Moving a candidate changes the coupling A built from query-key products with the diagonal removed and rows normalised. The figure shows A and its resolvent as Hinton squares (area is entry size, hollow is negative), the resolvent again as a 3D bar field, the Neumann hops stacked one per path length, and the scores before and after. Four small plots, one per ρ, compare the measured truncation error with the proved bound ρ^(H+1)/(1−ρ)·max|v|; the bound is right in shape and loose in size. Setting ρ to 1 prints the library's refusal. The queries and keys are toy functions of position; in the repository they are learned. On the repository's beds the multi-hop tail added nothing at K = 4 (resolvent minus one-hop = −0.0002 hit); the figure shows no trained result. Colour key: structure: the operator and its outputs, computed live; baseline: pointwise scores v, before any set interaction; proof: the Neumann tail bound and the 1/(1−ρ) inverse bound; seq: magnitude of a matrix entry. The stochastic form keeps the softmax and its diagonal. $P = \operatorname{softmax}(q k^\top/\sqrt d)$ is row-stochastic with $\rho(gP) = g$, so the pole sits at $g = 1$, and $$ O = (1 - g)\, P\, (I - gP)^{-1} V, \qquad g < 1, $$ is a convex combination of the rows of $V$. The implementation computes the softmax, zeroes NaN rows, and solves $(I - gP)$ against $V$ with a dense LU (`resolvent/resolvent.py:57-74`); $g \ge 1$ is refused with the message that $P$ is row-stochastic and the pole is at $g = 1$. **Figure 6.** The stochastic set resolvent O = (1−g)P(I−gP)^(−1)V on the reader's points. Four views at g = 0, 0.5, 0.9 and 0.99 show the weight matrix as bars over candidate pairs: at g = 0 it is one softmax read, and as g approaches 1 every row approaches the same weights. An inset shows the values V, their convex hull, and the outputs, which stay inside the hull at every g. A plot tracks the spread of the outputs across rows against g, and a cross marks the stationary mixture π·V that the outputs collapse towards. g = 1 is refused. The features are toy; this is the operator, not a trained ranker. Colour key: structure: the weight matrix and the outputs, computed live; proof: convex hull of the values; every output lies inside it; the cross is the stationary mixture π·V; baseline: g = 0, a single softmax read; seq: weight of one candidate on another. ### What a linear resolvent cannot do The first hypothesis for why set heads close the gap was that the gap is the multi-hop tail of the resolvent. It is not, and the reason is a small piece of algebra. The only linear permutation-equivariant coupling on $n$ candidates is $W = \alpha I + \beta J$, with $J$ the all-ones matrix. > **Claim: A linear equivariant resolvent is rank-inert.** > > Whenever $\rho(\gamma W) < 1$, > > $$ > (I - \gamma W)^{-1} s = \frac{1}{1 - \gamma\alpha}\left(s + \frac{\gamma\beta\, \mathbf 1^\top s}{1 - \gamma\alpha - \gamma\beta n}\, \mathbf 1\right), > $$ > > a positive rescaling plus a common shift, so the ranks of $s$ are unchanged. **Figure 7.** The reader sets α, β, γ and a score vector, and the real solve of (I − γ(αI + βJ))^(−1)s is plotted against s. Every output lies on the closed-form line of the rank-inertness equation, which rises with slope 1/(1 − γα) and in general has a nonzero intercept, so no two candidates swap. A slope graph sets the ranks of s against the ranks of the output. A 300-draws button replays the property test's sampling with a seeded generator in the browser; the count it reports inside the region is close to, not equal to, the repository's 139 of 300. Switching to a data-dependent coupling built from candidate features makes the ranks move. Panels where the spectral radius of γW reaches 1 are greyed: the library returns None there. Colour key: proof: the closed-form line: output as a function of input score; structure: numerically solved outputs; baseline: the pointwise scores s, the ranking before coupling; withdrawn: outside the region ρ(γW) below 1, where the library refuses. So a resolvent can change a ranking only through a data-dependent coupling $W(z)$, and that is what the set heads learn. Here is the head, in the form both set kinds share: ```python h = m = self.enc(z) if self.kind in ("hop1", "resolvent"): valid = mask[:, None, :] & ~torch.eye(z.shape[1], dtype=torch.bool, device=z.device) logit = (self.q(h) @ self.k(h).transpose(1, 2)) / h.shape[-1] ** 0.5 W = torch.softmax(logit.masked_fill(~valid, -1e9), -1) * valid # all-masked row -> zero row g = 0.99 * torch.sigmoid(self.theta) m = h + g * W @ h if self.kind == "hop1" else resolvent_apply(W, h, g) return self.out(torch.cat([h, m], -1)).squeeze(-1).masked_fill(~mask, -torch.inf) ``` `resolvent/heads.py:47-54 @ 3b1b7c1`. The one-hop head is $h + gWh$ and the resolvent head is $(I - gW)^{-1}h$, with one parameter set and $g = 0.99\,\sigma(\theta)$, so $g$ can never reach the pole. The library's `linear_equivariant_resolvent` returns `None` whenever the spectral radius of $\gamma W$ is at least $1 - 10^{-6}$, and a property test draws 300 random cases and checks that the ranks never move in the ones it can solve. ### LIN1 on the torus The third bed asks a different question: how few predictor calls does a good decision need? On the standard map, success is landing in a ball of radius $R$ around $G$ after $T$ steps. Linearising the rollout at the estimate and projecting on the unit goal direction $n_k$ gives a chance rule that costs one vector-Jacobian product per candidate: $$ \Pr(\text{success}_k) \approx \Phi\!\left(\frac{R - d_k}{\delta\, \lVert J_k^\top n_k \rVert}\right). $$ **Figure 8.** The Chirikov standard map on the 2-torus stretches a small start error until the posterior wraps the torus. For the selected candidate the figure shows 256 posterior members as dots and the tangent-linear ellipse that LIN1 reads from the Jacobian; scrubbing T shows where the ellipse stops describing the cloud. A table compares, per candidate, the Monte Carlo chance of success, the linearised chance and LIN1. Side panels show the predictor calls of the full rollout against the measured race, and the registered sign prediction for LIN1 against a three-member ensemble at long horizons, which was killed. One start is one realisation; it is not a ranking of the methods. Colour key: structure: posterior members, tangent ellipse, goal ball; baseline: latent-distance pick and full-rollout call count; proof: Monte Carlo posterior pick (the ceiling); measured: race call counts and recorded LIN1 − ENS3 gaps; withdrawn: the registered prediction that was killed; seq: probability of success. Across leads 12 to 32, LIN1's band-mean NS on seeds 3, 4 and 5 is 0.880, 0.881 and 0.879, against 0.729 to 0.737 for the best D-JEPA-spec operator on that bed. The Limitations section says what it ties with. ### The proofs underneath The Lean sources in the repository are the proofs of the CEQ lineage under `lean/CEQ/`, not proofs about the `resolvent/` library. They state, for the operators the set resolvent descends from, what the finite Neumann sum is and when it terminates. A read-only count script, `scripts/lean_count.py`, prints 13 files with 166 theorems and 41 lemmas, 207 in all. The pinned toolchain is Lean 4 v4.7.0. "0 sorry" below is a grep, not a compile: I did not build the proofs for this essay. Of the eight files in which `grep` finds the string, seven hits are the comment text "No `sorry`", and the eighth is a real `sorry` in a draft at `tests/foreman/phase_j/lean_draft_A5.lean`, outside the Lean build tree. No `sorry` appears as a tactic under `lean/`. The base is a telescoping identity, stated over any ring: **Lean 4 theorem occupancy_telescope** (0 sorry), [lean/CEQ/Occupancy.lean:52 @ 3b1b7c1](https://github.com/teerthsharma/resolvent/blob/3b1b7c1f52b848b81e97f01317cb397c624cf16c/lean/CEQ/Occupancy.lean#L52) ```lean theorem occupancy_telescope (A : R) (N : ℕ) : (1 - A) * occupancy A N = 1 - A ^ N := by ``` It says the truncated occupancy sum $\sum_{k 2 SE 0.039.. - Withdrawn: Hoeffding cascade on the torus (round 2, claim B). Killed by: cascade ratio 2.34–3.56, below 4; NS 0.9812 at seed 4, T = 32. The Bernstein race replaced it.. - Withdrawn: LIN1 beats a 3-member ensemble at long horizons. Killed by: round 2: LIN1 0.7589 against ENS3 0.7791 at T = 32; round 3: F killed at T = 40 (0.7543 against 0.7701); round 4: R0 killed at T = 40.. - Withdrawn: The verifier V2 null cannot be gamed. Killed by: plant r06 passed V2 and r07 passed V3 in round 3; fixed in round 4 by running the null in its own sandbox process.. **The beds are toys.** K = 4 at D = 2 on `shift`, K = 8 on a 2-torus, K = 63 at D = 8 with a known five-step map on `dial`. They were chosen because the truth is exact. The results say what a ranking operator can and cannot close given shared error, not how much shared error a trained JEPA has. Nothing here measures the shared fraction of a real world model, and no result is on a D-JEPA task. The D-JEPA operator is reimplemented from the paper's equations, not the released code, and has not been checked against the released checkpoints. **The resolvent is not separated from the alternatives.** The resolvent set head and the $\varepsilon = 4$ operator remain inseparable on the existing beds. The one-hop head matches the resolvent on bed `shift` to 0.0002 hit. The headline number at the top of this page belongs to a D-JEPA-spec Transformer operator with its bound widened until the bound is vacuous, not to the resolvent. **The $\sigma = 1$ regime is thin.** The 0.947 closure at $\sigma = 1$ is a share of a 0.024 gap, and that gap missed its own registered bar of 0.03. Per-seed closures range from 0.859 to 0.992 for the resolvent and above 1 for the one-hop head on one seed. At $\sigma = 10$ and 36 epochs, the resolvent head's top-1 agreement with the Bayes pick is 0.725, below the registered 0.750 gate (B11), even though its closure there is 0.958. The 12-epoch learning gate also failed (agreement 0.8865 against 0.90), which is why the closure bars were read at 36 epochs. **The rank-inertness test checks 139 of 300 draws.** The README calls it a 300-draw property test. The test draws 300 cases, but only the ones inside $\rho(\gamma W) < 1$ can be solved; its own comment says 139 of the 300 fall inside, and it asserts that at least 100 were checked. The result rests on the algebra; the test checks it on 139 cases. **LIN1 ties the ensemble.** On seeds 3 to 5 at leads 12 to 32, LIN1's band-mean NS is 0.8799, 0.8806 and 0.8794 and the three-member ensemble ENS3's is 0.8796, 0.8808 and 0.8802: within 0.001 NS. The comparison in the README is against the best D-JEPA-spec operator, which LIN1 does beat. The torus verdict file records `lin_beats_ens3` false, `cert_never_hurts` false and `topo_pass` false, and the README reports none of these flags. The B2 race's configuration rests on two tuning seeds, and its tail saving at T = 32 is not bound. **The resolvent costs more than DeepSets.** The cost bench registered "fastest" as a claim that needs matched quality and per-decision wall-clock at or below every matched comparator. The quality side holds: mean regret over three seeds is 0.0196 for the resolvent against 0.01879 for DeepSets and 0.01802 for the pointwise head, with the latent-distance floor at 0.12153, so the resolvent is within 0.0008 of DeepSets, inside the registered 0.01 margin. The speed side does not. At batch 64, the dense-LU resolvent takes a median 3,203.2 µs at K = 64 against 1,212.3 µs for DeepSets and 1,364.7 µs for the set transformer, and 61,897.7 µs at K = 1,024 against 2,502.7 µs for DeepSets. The four-hop Neumann variant is faster than the dense solve (24,031.2 µs at K = 1,024) and still slower than DeepSets. No verdict was written for the cost bench, and the README does not report it. **A null band was replaced after the data came in.** Round 2 registered that a head trained on shuffled labels should score within ±0.05 NS of zero. It did not: the shuffled-label head hit 0.0300 against a chance rate of 0.0159, 5.0 standard errors above it, so chance was the wrong null. The criterion was replaced by "null hit at most distance hit plus 3 SE". The recorded null NS at $f = 1$, $\sigma/\rho = 3$ is 0.089 to 0.132, outside the original band. At the first registered evaluation size of 2,000 starts the bed could not resolve the 0.01 and 0.05 NS bars, since one start moved NS by 0.033, and the evaluation size was raised to 20,000. **The edge law is a fit.** After P2n and P2r died, round 4 registered an affine law for the edge of the $\varepsilon = 4$ operator over D-JEPA as specified: $$ \mathrm{edge} = -0.0683 + 0.6618 \times \mathrm{closable}, $$ with closable defined as one minus the distance arm's NS and the edge as $\mathrm{NS}(\texttt{dj4L}) - \mathrm{NS}(\texttt{dj02})$. It was fitted on 15 cell-seeds after reading the data, and then predicted a fresh cell, (0.85, 2), at 0.037, 0.063 and 0.059 against measured 0.035, 0.060 and 0.062, within 0.004 against two standard errors of 0.036 to 0.039. In the README's words: "the law is a fit that survived one test, not a derivation." It has been tested on one cell. **Figure 10.** How the edge law was found. Each point is one cell-seed of bed dial: x is the room left to close (one minus the distance arm's NS), y is the edge of the ε = 4 operator over D-JEPA as specified. Two pre-registered laws, a threshold at 0.25 and a line through the origin, are drawn dashed; both were killed. The solid line is the affine law fitted afterwards on 15 cell-seeds; the three diamonds are the fresh cell it then predicted. Whiskers are one standard error. Hollow rings are nine re-scored cells outside the fit, where the law is not claimed; the dashed grey line is the registered bar, 0.05. Rounds 3 and 4 differ in σ/ρ and evaluation size, and the law pools them. Colour key: measured: cell-seeds the law was fitted on, and the fresh cell it predicted; structure: the fitted affine law; ink-3: context cells outside the fit, where the law is not claimed; withdrawn: pre-registered laws that were killed. **The verifier is not an OS boundary.** Its sandbox is an in-process audit hook, and a candidate can still detect its null through the correlation between labels and distance. **The Lean does not prove the library.** The 207 declarations belong to the CEQ lineage. No declaration names `set_resolvent`, `build_A`, the relational operator or the candidate set. The signed coupling falls outside `PerronCertificate`, which needs nonnegative entries. Permutation equivariance, the rank-inertness result and the hull limit are proved in prose and checked by tests, not in Lean. I did not compile the proofs for this essay, and "0 sorry" is a grep. ## Read more [View the project](https://github.com/teerthsharma/resolvent) · [Source on GitHub](https://github.com/teerthsharma/resolvent) - The repository: [github.com/teerthsharma/resolvent](https://github.com/teerthsharma/resolvent), read at commit `3b1b7c1f52b848b81e97f01317cb397c624cf16c`. - The project site: [teerth.dev/resolvent](https://teerth.dev/resolvent/). The site's pages still tell the CEQ family's story; they are regrouped, not rewritten. - The [README](https://github.com/teerthsharma/resolvent/blob/3b1b7c1f52b848b81e97f01317cb397c624cf16c/README.md), including the "What we got wrong" table and the limits. - The library: [`resolvent/`](https://github.com/teerthsharma/resolvent/tree/3b1b7c1f52b848b81e97f01317cb397c624cf16c/resolvent), with `resolvent.py`, `heads.py`, `djepa.py` and the three beds. - The Lean proofs: [`lean/CEQ/`](https://github.com/teerthsharma/resolvent/tree/3b1b7c1f52b848b81e97f01317cb397c624cf16c/lean/CEQ), in particular `Occupancy.lean`, `Nilpotent.lean`, `Contraction.lean` and `OracleSeparation.lean`. - The results and their pre-registrations: [`experiments/`](https://github.com/teerthsharma/resolvent/tree/3b1b7c1f52b848b81e97f01317cb397c624cf16c/experiments), one `BAR.md` per round; the verifier in [`daedalus/`](https://github.com/teerthsharma/resolvent/tree/3b1b7c1f52b848b81e97f01317cb397c624cf16c/daedalus). - The failure ledgers: [`docs/FAILS.md`](https://github.com/teerthsharma/resolvent/blob/3b1b7c1f52b848b81e97f01317cb397c624cf16c/docs/FAILS.md), [`MISTAKES.md`](https://github.com/teerthsharma/resolvent/blob/3b1b7c1f52b848b81e97f01317cb397c624cf16c/MISTAKES.md) and the retraction row C10 in [`docs/canon/CORRECTIONS.md`](https://github.com/teerthsharma/resolvent/blob/3b1b7c1f52b848b81e97f01317cb397c624cf16c/docs/canon/CORRECTIONS.md). - The work it builds on: [D-JEPA, arXiv 2609.24749](https://arxiv.org/abs/2609.24749). - Related essays on this blog: [caustic](/caustic), whose Theorem 1 (`orbit_error_bound`) is proved in the same Lean tree, and [Epsilon-Hollow](/epsilon-hollow), whose finite-statement convention the occupancy proofs follow. --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # sigmoid > Measuring the shape of a trained model's activations as a state, and rolling that state forward with one ridge solve instead of a second network. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/sigmoid - Repository: https://github.com/teerthsharma/sigmoid - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "sigmoid", 2026, https://teerth.blog/sigmoid - Languages: Python - Question: Can the persistent homology of an activation window be a state that a closed-form operator rolls forward, and can the rollout say when to stop being trusted? - Headline result: 0.855 island partition read from ψ, probe accuracy on the S²-Rips corpus (control: PCA channel 0.469 · majority class 0.487) ## What it is A trained network already computes a rich internal state. What it does not give you is a dynamics over that state that can be evaluated without paying for the network again. Every forward pass rebuilds the state from scratch, and nothing in the architecture exposes "what happens next" as an object one can iterate cheaply. The usual answer in world-model research is a second network: a latent dynamics model fit by gradient descent, as in the Dreamer family's recurrent state-space models or in diffusion world models. That answer brings training cost, hyperparameters, and rollouts that drift without warning, because nothing constrains them. sigmoid is my attempt at the other route. It takes any producer of hidden activations (a transformer, a policy network, a simulator, a bare callable) and builds a world model around it without touching its weights. Three choices define it. The state is not learned; it is measured, as the persistent homology of a sliding window of activations. The dynamics is not trained; it is solved in closed form, by one ridge regression. And the trust placed in a rollout is bounded by the Banach fixed-point theorem where the fitted operator contracts, and checked by a self-consistency gate where it does not. I want to state the result before the machinery, because the repository states it plainly and the rest of this essay should be read against it. sigmoid **represents** entity-structured topology that a linear channel of the same data does not recover. It does **not yet predict** with it. On an entity corpus on the sphere, the topological channel ψ reads the island partition at 0.855 probe accuracy against 0.469 for the PCA channel; on a granular contact simulation it reads the number of contact components at 0.778 against 0.104. Those are representation results. On prediction the record is negative: on the corrected Lorenz table, topology helps at short horizon and hurts at sixteen steps (normalized RMSE 1.1266 against 0.9889 for the same linear channel without ψ), and on the residual stream of distilgpt2 sigmoid does not beat predicting the dataset mean at long horizon. The repository's own summary names the gap between those two verbs as the remaining research programme, and I agree with it. What does hold is cost. One imagined step is a matrix-vector product: 97 µs against 86 ms for a real distilgpt2 forward pass on 214 tokens, a ratio of about 880. That compares cost only, not the quality of the imagined state. **Figure 1.** The pipeline the repo defines: a model's activations are captured, encoded to z = \[ψ ; u], advanced by the fitted operator T, and read by the gate. Each stage carries the repo's measured cost on its RTX 4060 laptop host (forward pass 86 ms, encode one window 0.909 ms, imagine one step 97 µs, gate read 11–41 µs). The reader can time a stand-in matrix-vector step and a toy encode on their own device; that stand-in uses a random operator and says nothing about the fitted one. The ratio compares cost only, not the quality of the imagined state. Colour key: baseline: the forward pass the operator replaces in the loop; structure: a stage of the sigmoid pipeline, at its repo-measured cost; measured: timed on this device, stand-in operator; parameter: time axis and number of imagined steps k. Every number here is the repository's own, measured on its host (Python 3.11.9, numpy 1.26.4, scipy 1.17.1, torch 2.5.1, an RTX 4060 laptop GPU) at commit `71d8a21`, package version 0.1.0. I did not re-run any of it for this essay. ## What it can do ### Read a partition that a linear channel misses The cleanest positive result is on a synthetic entity system: points moving on the unit sphere, where two entities are joined when their geodesic distance is below a fixed radius, and the "islands" are the connected components of that graph. **Measured: 0.855** island partition read from ψ (spatial cloud, absolute radii), S²-Rips corpus. Control: PCA channel u 0.469 · majority class 0.487. Source: BENCHMARKS.md:26-29 @ 71d8a21 · RTX 4060 laptop host, repo-reported. The same table carries an ablation that matters more to me than the headline. With a temporal cloud (successive state vectors as the points) the probe reads 0.386, below the 0.487 majority floor. With the spatial cloud (each observation reshaped into its entities) it reads 0.764. With the spatial cloud and absolute radii it reads 0.855. Only two configuration choices differ; they are the subject of the section on what is new. The partition can also be read ahead of time on this system. Probing ψ at lag 0, 1, 4 and 16 gives 0.920, 0.817, 0.626 and 0.494, against 0.311, 0.308, 0.305 and 0.201 for the linear channel. On a second system the forecasting story does not survive, and I show that next to the representation result rather than leaving it for later. **Measured: 0.778** contact components read from ψ, granular contact (24 soft disks, 1999 steps, 17 distinct counts). Control: 52-dim linear channel 0.104 · majority 0.200 · persistence baseline 1.000 at lag 0. Source: BENCHMARKS.md:28-29; SIGMOID.md:570-585 @ 71d8a21 · repo-reported. At lag 1 the same probe reads 0.265 against a persistence baseline (predict that the component count stays as it is) of 0.250. That is barely better than copying the present. The persistence baseline was not in the original specification; the repo notes that without it, lag 1 would have looked like successful forecasting. **Figure 2.** Entities move on the sphere by geodesic steps; an edge joins two entities within geodesic radius 0.95, and the islands are the components of that graph. The timeline compares the flood-fill island count with β₀ read at the absolute radius and with β₀ read after dividing by the diameter, on the same frames. Below, the repo's measured probe accuracies: temporal 0.386, spatial 0.764, spatial with absolute radii 0.855, against linear 0.469 and majority 0.487, and the forecast lags for S²-Rips and for granular contact, where the persistence baseline is drawn and cells where ψ does not beat it are marked. The browser corpus uses a different random stream from the repo's; the accuracies are the repo's, not recomputed. Colour key: structure: β₀ read at the absolute radius; ψ probe accuracy; baseline: β₀ read with the diameter divided out; persistence baseline; measured: flood-fill ground truth, linear channel and majority floor; withdrawn: a forecast cell where ψ does not beat persistence; parameter: the frame cursor on the timeline. ### Imagine cheaply, and check without the model **97 µs** imagine one step **0.909 ms** encode one window, 768-dim, W = 24 **11–41 µs** gate read per window The split is deliberate: expensive topology at calibration, arithmetic at runtime. Persistent homology runs only when a window of real activations is encoded; an imagined step is a matrix-vector product, and the gate that checks it is a few norms with no call to the wrapped model. ### Build the same sparse-attention schedule as the merged kernel, faster The merged Triton kernel in [`triton-lang/kernels#22`](https://github.com/triton-lang/kernels/pull/22), which came out of my Aether-Lang work, builds a causal block schedule for sparse attention from a zero-dimensional persistence salience over key-block centroids: sink blocks, a local window, and the most salient blocks. The salience of a block is the single-linkage merge height at which its component was absorbed, which is the same H₀ object sigmoid computes for its state. `sigmoid.triton` builds that schedule from the minimum spanning tree instead of from all pairs. **Measured: 11.6–15.0×** schedule builder against the merged kernels#22 reference, 64 to 512 key blocks (CPU builders, triton stubbed). Control: merged reference builder: 7.02 ms at 64 blocks, 637.23 ms at 512; sigmoid 0.60 ms and 47.28 ms. Interval: max abs salience deviation 1.8e-15 to 3.6e-15; full CSR schedules bit-identical at sequence 512, 1024, 2048. Source: BENCHMARKS.md:121-132 @ 71d8a21 · RTX 4060 laptop host, repo-reported. For autoregressive decode, `IncrementalSalience` grows the tree one key block at a time instead of rebuilding it. The repo measures 8.6× to 21.7× per appended block between 64 and 512 blocks, bit-exact across 2748 append-against-rebuild comparisons over 200 tie-heavy configurations. None of that makes sparse attention pay on the model this repo can run. End to end on distilgpt2, best of three, dense attention generated 102, 97 and 95 tokens per second at contexts 256, 512 and 960; the topology schedule generated 51, 70 and 78. The schedule builder is fast; the attention it schedules is slower here, and the details are in the limitations. ### Run inference with a schedule, and measure what it costs The inference path is checked against a control first: its dense backend reproduces HuggingFace `generate(do_sample=False)` token for token, 64 of 64, and a saturated schedule through the sparse path stays within KL 1e-9 of dense. With the schedule actually sparsifying, at context 960 over 48 greedy tokens: **Measured: 0.11923 nats** KL against dense, topology schedule top-k 4, local radius 1, density 0.481. Control: same path with top-k 0 (no topology), radius 1, density 0.350: KL 0.99315. Interval: top-k 8, radius 2, density 0.721: KL 0.03960; greedy tokens differing from dense: 46/48 (top-k 8), 46/48 (top-k 4), 47/48 (top-k 0). Source: BENCHMARKS.md:315-325 @ 71d8a21 · repo-reported. The densities are not equal, so this is not a matched comparison; it says that removing the topology blocks entirely, at a density not far below, costs about eight times the KL. Nearly every greedy token still differs from dense in all three sparse rows, so none of them is a drop-in for dense decoding. ### Plan, and refuse The operator is action-conditioned, so it can be planned against. `TopologicalMPC` draws Gaussian action sequences, rolls each through the operator, rejects any candidate whose worst gate score reaches 1.0, and refits the sampling distribution to the cheapest survivors. If none survive, the plan reports `feasible = False` and the robot holds still. A world model asked to drive a robot should be able to say it does not know; this is how sigmoid says it. Whether the topological cost term helps that robot is a separate question, answered negatively in the limitations. On images, the Euler characteristic of a 4-connected binary mask, $\chi = F - H - V + Q$, is four array reductions: 0.044 ms on a mask of 5025 foreground pixels, against 127 ms for Rips H₀ on 419 subsampled pixels of the same image. The two sides compute different objects on different inputs, so this is a cost of reading, not a speedup of one algorithm over another. ## How it was made ### The window, and why it is standardized Given a hidden trajectory $h_1,\dots,h_T\in\mathbb{R}^D$ and a window length $W$, the point cloud at time $t$ is $X_t=\{h_{t-W+1},\dots,h_t\}$. Before any distance is taken, each coordinate is centred by its median and scaled by $1.4826\cdot\mathrm{MAD}$, the factor that makes the median absolute deviation match a Gaussian standard deviation, with fallbacks when the scale is at or below $10^{-6}$. This step is not cosmetic. On distilgpt2 the largest per-dimension standard deviation was 22.1 against a median of 0.308, a 72× ratio concentrated in about five channels; without robust scaling the whole barcode would be a function of those five channels. ### H₀ without a simplicial complex > **Definition: H₀ barcode of a window.** > > For a Vietoris–Rips filtration, the death times of zero-dimensional classes are the single-linkage merge heights, which are the edge weights of the Euclidean minimum spanning tree. Every class is born at 0, and one class never dies. $$ B_0(X)=\{(0,w):\,w\in\mathrm{MST}(X)\}\ \cup\ \{(0,\infty)\} $$ **Figure 3.** The cloud lies on the floor and the vertical axis is the Rips radius. The single-linkage tree stands over the cloud: each merge happens at the length of a minimum spanning tree edge, so the order in which a rising disc cuts the tree is the barcode. The component count at the chosen radius is computed twice, from the bars and from an independent union-find on the thresholded graph, and the two are checked against each other live. The grey band on the β₀(r) curve is the log-height plateau of this frame. Presets include tied lattices and coincident points, the inputs that broke earlier versions. H₀ only; a barcode cannot be inverted back to coordinates. Colour key: structure: the single-linkage tree and the H₀ barcode; proof: bars and union-find agree on β₀; parameter: the filtration radius the reader moves. The implementation is short, and two of its comments record bugs I had to find before the numbers in this essay meant anything. ```python # sigmoid/state.py:100-134 @ 71d8a21 (docstring trimmed) def h0_barcode(points: np.ndarray) -> Barcode: pts = np.asarray(points, dtype=np.float64) if pts.ndim != 2 or pts.shape[0] < 2: return Barcode(dim=0, bars=np.array([[0.0, np.inf]]), diameter=0.0) dist = _pairwise(pts) diameter = float(dist.max()) if diameter <= 1e-12: return Barcode(dim=0, bars=np.array([[0.0, np.inf]]), diameter=0.0) dist = dist / diameter from scipy.sparse.csgraph import csgraph_from_dense, minimum_spanning_tree # csr treats an explicit 0.0 as *no edge*, so coincident points lose the # edge between them and the tree routes around it: [[0,0],[0,0],[1,0]] came # out with merge heights [1, 1] instead of [0, 1], inventing a component # that separates at O(1). Same failure the Gram identity caused in # `_pairwise`, one layer down. null_value=inf keeps zero edges; they are # indistinguishable from absences on the way back out, so restore them by # count -- a spanning tree on n points has exactly n-1 edges. graph = csgraph_from_dense(dist, null_value=np.inf) mst = minimum_spanning_tree(graph).toarray() heights = np.sort(mst[mst > 0.0]) heights = np.concatenate([np.zeros(pts.shape[0] - 1 - heights.shape[0]), heights]) bars = np.empty((heights.shape[0] + 1, 2), dtype=np.float64) bars[:-1, 0] = 0.0 bars[:-1, 1] = heights bars[-1] = (0.0, np.inf) # the essential class return Barcode(dim=0, bars=bars, diameter=diameter) ``` The other bug lives in `_pairwise`. The Gram identity $\|x\|^2+\|y\|^2-2x\cdot y$ is one matrix multiply, and it was used until it was caught cancelling to exactly 0.0 for points 1e-12 apart, which merged distinct points inside the barcode. The distances now come from `cdist`, which subtracts before squaring, and the 0.909 ms encode cost above already includes that choice. ### From a barcode to a vector a matrix can act on A barcode has no fixed length, so no linear operator can act on it directly. The numerator of its Hilbert series supplies one: births count positively, deaths negatively, and binning the exponents gives a vector of fixed length. $$ N(s)=\sum_i s^{b_i}-\sum_i s^{d_i},\qquad c_k=\#\{\text{births in bin }k\}-\#\{\text{deaths in bin }k\},\quad k=\lfloor b\cdot\mathrm{degree}\rfloor $$ **Figure 4.** A barcode with any number of bars becomes D numbers: N(s) is drawn on \[0, 1], its binned coefficients c_k sit below it (hatched bars are negative), and the full ψ vector is shown as four segments (Hilbert coefficients, Betti samples at normalized radii, counts above absolute radii, four geometry scalars), 44 coordinates at the default configuration. A sweep moves two clusters apart and stacks c(g) as columns, so the reader sees deaths migrate across bins. The readout shows the true binned sum, including the case where a death at exactly 1.0 falls out of the last bin. Colour key: structure: the barcode and its Hilbert coefficients; parameter: the cluster gap g the reader moves. ```python # sigmoid/state.py:168-186 @ 71d8a21 def hilbert_coefficients(barcode: Barcode, degree: int) -> np.ndarray: """Hilbert-series numerator coefficients of a barcode. N(s) = sum_i s^{b_i} - sum_i s^{d_i} Coefficient k counts births in bin k minus deaths in bin k. The whole barcode collapses to a fixed-length vector this way, which is what makes a linear operator on topology possible at all. """ coeffs = np.zeros(degree, dtype=np.float64) for birth, death in barcode.bars: ib = int(np.floor(birth * degree)) if 0 <= ib < degree: coeffs[ib] += 1.0 if np.isfinite(death): idd = int(np.floor(death * degree)) if 0 <= idd < degree: coeffs[idd] -= 1.0 return coeffs ``` The name is grander than the computation. Nothing algebraic is computed; $c$ is a signed histogram of birth and death heights at resolution $1/\mathrm{degree}$. The repo states an identity, $N(1)=\sum_k c_k$ equals the number of essential classes, and $N(1)$ itself does equal it. The binned sum does not always: a finite death that normalizes to exactly 1.0 lands in bin `degree`, which the loop excludes, and the test for the identity uses hand-made deaths of 0.4 and 0.7 only. With the default configuration, $\psi$ concatenates 24 Hilbert coefficients, Betti numbers at 8 normalized radii, counts of deaths above calibrated absolute radii, and four geometry scalars, 44 coordinates in all. The full state is hybrid, $z=[\psi;u]$. ψ carries regime: how many components, how spread, how confined. $u$ carries content: a whitened PCA projection of the standardized observation, the part that decodes back to activations. Neither half is enough alone. ψ cannot emit predicted activations, because topology forgets coordinates. $u$ alone is a plain linear autoencoder, which is exactly the null model a topological claim has to beat, and keeping it as a separate channel is what lets every result in this essay be stated against it. ### The operator is solved, not trained $$ W=\arg\min_W \sum_t\big\|W\varphi(z_t,a_t)-z_{t+1}\big\|^2+\lambda\|W\|_F^2=(X^\top X+\lambda I)^{-1}X^\top Y,\qquad \varphi(z,a)=[\,z;\ a\otimes z;\ a;\ 1\,] $$ **Figure 5.** The repo's action-conditioned ridge operator fitted on a small system the reader sets: the lift \[z ; a⊗z ; a ; 1], the closed-form solve, and the fitted entries against the true ones, with 16-step rollouts at two actions. The 3D panel is the loss restricted to two entries of W; gradient descent from a draggable start walks down a convex bowl to the point the closed form lands on directly. Three λ panels show shrinkage. The system is linear by construction; the repo's real systems are not, and there the operator is judged only on held-out rollouts. Colour key: measured: the true system; structure: the ridge-fitted operator; loss height on the bowl; baseline: gradient descent, the iteration the closed form avoids; proof: closed form and descent limit agree. The lift makes the operator bilinear in state and action: $W\varphi = T_0 z+\sum_k a_k T_k z + Ba + c$. The default ridge is $\lambda=10^{-3}$. The fit is the four lines below; there is no learning rate and no epoch. ```python # sigmoid/operator.py:332-341 @ 71d8a21 X = self._lift(X_states, actions) if self.block_split and 0 < self.block_split < self.state_dim: self.W_ = self._fit_block_diagonal(X_states, Y, actions) else: gram = X.T @ X + self.ridge * np.eye(X.shape[1]) self.W_ = np.linalg.solve(gram, X.T @ Y).T self._project_spectral() residual = Y - X @ self.W_.T self.step_rmse_ = float(np.sqrt(np.mean(residual**2))) ``` Whether ψ and $u$ should evolve jointly or as two independent blocks is chosen from data rather than assumed. `SigmoidWorldModel` splits the training trajectories chronologically 80/20, fits both structures on the first part, rolls both forward on the held-out part over a horizon of up to 16 steps, scores the error on the $u$ channel only, and keeps the winner. A convex fit does not make the model accurate. The fitted operator is reported with its measured $\rho$; the spectral clip runs only when `rho_max` is set, and its default is `None`. ### What a rollout is worth: the certificate and the estimate Let $A$ be the state-to-state block and $\rho=\sigma_{\max}(A)$, with $\varepsilon$ the one-step residual. If $\rho<1$, the rollout map contracts and the accumulated error is bounded by a geometric series. Propagating the residual covariance instead gives a sharper number that is an estimate, not a bound. $$ E(n)\ \le\ \varepsilon\,\frac{1-\rho^n}{1-\rho}\ \xrightarrow[n\to\infty]{}\ \frac{\varepsilon}{1-\rho},\qquad C_n=A\,C_{n-1}A^\top+\Sigma,\quad C_0=0,\quad \text{reported as }\sqrt{\operatorname{tr}(C_n)/d} $$ **Figure 6.** On the test suite's synthetic system (ten dimensions, anisotropic residuals), the scalar Banach bound, the covariance-propagated estimate and the measured k-step error are drawn on one log axis, and in the error space as a sphere (the scalar bound) and an ellipsoid (the estimate's top three directions) around the measured error vectors. A residual-autocorrelation readout warns when the estimate stops being trustworthy. The repo's distilgpt2 certificate table, where the scalar bound is vacuous at every setting, is printed verbatim beside it. Colour key: baseline: the worst-case scalar bound; structure: the covariance-propagated estimate; measured: measured rollout error, its 5th to 95th percentile band, and the held-out error vectors; parameter: horizon k, tolerance τ, and the safe horizon where each curve meets τ; proof: scalar bound at or above measured error at every step from k = 2 to 32. The scalar bound is true and, on real data, useless. On distilgpt2 the fitted operator has $\rho=1.2556$ and the bound at horizon 16 is 368. Clipping $\rho$ to 0.99 or 0.90 brings it to 12.5 or 6.95, still above 1.0, so never informative, while one-step NRMSE gets worse with each clip (0.2410, 0.2576, 0.2602). **Measured: 12.6×** covariance estimate tighter than the scalar bound, anisotropic synthetic system (ρ 0.983). Control: scalar bound 5.065; estimate 0.402; measured error 0.401 (estimate/measured 1.00, residual autocorrelation 0.000). Interval: AR(1) residuals: estimate/measured 0.42 (autocorr 0.808); Lorenz: 0.22 (autocorr 0.937). Source: BENCHMARKS.md:172-182 @ 71d8a21 · repo-reported. The estimate assumes zero-mean residuals that are uncorrelated across steps, and when that fails it under-bounds: 2.4× low with AR(1) residuals, 4.6× low on Lorenz. The fitted operator stores its residual lag-1 autocorrelation so that the reader of a certificate can tell which regime they are in. An early version clipped $\rho$ to 0.995 so that a certificate always existed; on Lorenz that raised one-step NRMSE from 0.067 to 0.317. Chaotic systems have $\rho>1$ by definition, so the clip bought a presentable certificate by misreporting the dynamics. Contraction is now reported, never forced. ### Stabilizing a rollout without rewriting the operator When a rollout must not blow up, there is a second way to get there that leaves the fitted operator alone: a proportional-derivative correction toward the fixed point $z^*$, refused when its gains violate a stated condition. $$ u_t=-\Big(\alpha e_t+\beta\,\frac{e_t-e_{t-1}}{dt}\Big),\qquad e_t=z_t-z^*,\qquad z_{t+1}=T(z_t)+dt\,u_t,\qquad \alpha+\frac{\beta}{dt}<1,\ dt\ge 1 $$ **Figure 7.** One expansive two-dimensional operator rolled out three ways: plain, spectrally clipped, and with the PD correction, drawn as spirals rising in time. The per-step energy-descent check runs live, and the gain condition is enforced as the code enforces it: when α + β/dt is not below 1 the stabilized rollout is refused. The stability panels over (α, β) are computed by the figure from the closed loop the code builds; they are the figure's own derivation, not a repo result. Colour key: baseline: plain and clipped rollouts; structure: the PD-corrected rollout; proof: energy descent holds at this step; gains with closed-loop radius below 1; withdrawn: energy descent violated at this step; unstable gains (hatched); parameter: the step axis and the gains α, β the reader moves. **Measured: 1.49 → 3.30e-03** state norm over 40 steps with the PD correction (α 0.5, β 0.2, dt 1), expansive operator ρ = 1.600. Control: plain rollout of the same operator: 2.98 → 2.72e+08. Source: BENCHMARKS.md:480-485 @ 71d8a21 · repo-reported. The trade is explicit in the code: the correction biases the rollout toward the fixed point, which is right when a long rollout must not blow up and wrong when the true trajectory of an expansive system is what you need. The docstring attributes the gain condition to `AetherGovernor.lean`, a proof that is not in this repository, so nothing in this repository checks it. ### The gate: checking an imagined state against itself An imagined state cannot be checked against the truth, because the truth is the computation I declined to run. It can be checked against itself. ψ and $u$ are two local views of one activation window, so on real data a map $R$ from one to the other can be fitted by ridge on $[u;1]$, and on a genuine state the views agree. Nothing in the operator's fit forces an imagined $\hat\psi$ to stay the signature of an imagined $\hat u$, so their residual grows when a rollout leaves its calibrated region. A third term reads the window's effective rank, abnormally low under repetition and high under noise. $$ r=\big\|R\,\hat u-\hat\psi\big\|,\qquad \mathrm{er}=\exp\Big(-\sum_i p_i\log p_i\Big),\quad p_i=\frac{s_i^2}{\sum_j s_j^2} $$ **Figure 8.** The effective-rank term computed live on synthetic windows (smooth, repeated token, noise, and a mix): the window's rows in their top three singular directions, the calibration histogram of log-rank deviation with its 0.98-quantile threshold, and the score that fires at 1.0. The centred and uncentred ranks are shown side by side, because centring hides the repeated-token case. Beside it, the repo's measured out-of-distribution scoreboard, in which the gate does not win. The sheaf residual r needs a fitted R on language-model states and appears only as measured numbers. Colour key: structure: calibration distribution and window spectrum; baseline: the fixed threshold; detectors that need a forward pass; measured: measured AUROC of each detector; parameter: window mix and quantile the reader moves. ```python # sigmoid/sheaf.py:207-212 @ 71d8a21 predicted_psi = self.R_ @ np.concatenate([u, [1.0]]) sheaf_residual = float(np.linalg.norm(psi - predicted_psi)) manifold_distance = float(np.linalg.norm((z - self.mean_) * self.inv_std_)) sheaf_score = sheaf_residual / self.sheaf_threshold_ manifold_score = manifold_distance / self.manifold_threshold_ ``` Each term is divided by its 0.98-quantile on calibration data, the gate's score is the maximum of the three, and it fires at 1.0, so a score reads as "how unusual this would have been during calibration". The support term, as coded, scales each coordinate by its own standard deviation; it is not a covariance-whitened Mahalanobis distance, though the docs call it one. **Measured: 0.846** mean AUROC of the gate over five out-of-distribution types, 332 held-out prose windows. Control: Mahalanobis 0.899 · PCA reconstruction 0.888 · sheaf term alone 0.869 · kNN 0.836. Interval: read cost per window: gate 11–41 µs, PCA 14–54 µs, Mahalanobis 160–395 µs. Source: BENCHMARKS.md:186-211 @ 71d8a21 · repo-reported. As a detector of unusual input, the gate loses to cheaper baselines. What survives is applicability: it is the one entry in that table that can read an imagined state, because the others need either a real window or a forward pass. ## What's new in it **Against a learned latent dynamics model.** Dreamer-style recurrent state-space models and diffusion world models learn the state and the dynamics by gradient descent and offer no rollout guarantee. sigmoid measures the state, solves for the dynamics in one convex step (seconds for 8226 transitions), and attaches a bound and a gate to every rollout. The price is linearity: a learned recurrent or diffusion model is strictly more expressive and will beat it on dynamics that genuinely need nonlinearity; the repo's own example is distilgpt2, where sigmoid does not beat the mean at long horizon. **Against the usual persistent-homology pipeline.** Most uses of topological features in machine learning compute a barcode and vectorize it, without stating two choices that sit upstream of every such pipeline. Which cloud: the window of successive states, which describes the shape of a trajectory segment, or each observation reshaped into its entities, which describes their arrangement at one instant. Which scale: whether filtration values are divided by the diameter, which buys scale invariance by discarding scale. On the S²-Rips corpus those two binary choices moved the probe from 0.386, below the majority floor, to 0.855. Neither wrong choice raises an exception; both produce a ψ with healthy variance and no information. sigmoid makes both explicit configuration (`entity_dim`, `abs_radii`) and carries a selector that picks the cloud without labels. The repo also states a hypothesis I hold but have not tested: that many published negative results for topological features are specification errors of this kind. The agenda file lists the audit that would test it. **Against reporting the win.** Every rollout comparison in the repo is made at a matched budget against a `no_topology_same_u` arm: the same linear channel with ψ deleted. That arm exists because an earlier comparison against a `linear_only` arm raised the PCA rank along with the topology, and the repo's first Lorenz headline ("\~30% better") turned out to be rank, not topology. The bench reports failed gates rather than dropping the losing arms. **A condition, not a slogan.** The repo ends with a two-clause statement of when topology pays, and the second clause is what the negative results taught me. Representation needs a band of radii, fixed relative to the configuration's own scale, on which H₀ is constant; the width of that band is how much error in stating the radius the encoder absorbs. Prediction helps only when the partition is an input to the dynamics, not merely an observable of them; a partition that is only an observable carries no force on the future. The first clause replaced an earlier "fixed physical distance" clause after a control falsified it, as the limitations below record. ## What no one else built The parts of sigmoid are not new; here is where each comes from, and what differs. **Topology of activations.** Persistent homology of a network's internal representations is an established line of work. [Naitzat, Zhitnikov and Lim (JMLR 2020)](https://jmlr.org/papers/v21/20-345.html) track Betti numbers of a dataset as it passes through the layers of a trained classifier. [Gardinazzi et al. (ICML 2025)](https://arxiv.org/abs/2410.11042) use zigzag persistence to follow how features of prompt representations appear and vanish from layer to layer in large language models, and use it as a pruning criterion. [Kushnareva et al. (EMNLP 2021)](https://arxiv.org/abs/2109.04825) build topological features of attention maps to detect machine-generated text. In each of these, topology is a descriptor: of depth, of a dataset, or a feature for a classifier. The difference in sigmoid is concrete: a vectorized H₀ barcode of a sliding time window is used as the state of a dynamics model, which is fitted, rolled forward and checked. That is also where sigmoid fails, which the prior work does not attempt and so cannot fail at. **The operator.** Fitting a linear operator by least squares on a dictionary of observables is extended dynamic mode decomposition ([Williams, Kevrekidis and Rowley, 2015](https://arxiv.org/abs/1408.4408)), and the bilinear lift $[z;\,a\otimes z;\,a;\,1]$ is the bilinear Koopman realization whose advantages for control [Bruder, Fu and Vasudevan (2021)](https://arxiv.org/abs/2010.09961) analyse. I do not claim the operator. The difference is the dictionary and the reporting: the observables are Hilbert coefficients and Betti counts computed from another model's activation window, beside a PCA channel that serves as the built-in null model, and the spectral norm of the fit is reported rather than clipped. **The gate.** Measuring how far local data fail to glue is [Robinson's consistency radius](https://arxiv.org/abs/1603.01446) for sheaves of sensor data, where the restriction maps come from the sensor model and the stalks hold measurements. In sigmoid there is one restriction map, learned by ridge between two channels of the same window, and it is evaluated on imagined states where no measurement exists. That changes what the residual can be used for, not how well it detects: as a detector of unusual input it loses to the [Mahalanobis-distance score](https://arxiv.org/abs/1807.03888) and to PCA reconstruction. What I claim for it is the property that the measured scoreboard leaves standing, that it reads an imagined state without a model call. **The schedule.** That single-linkage clustering is determined by the minimum spanning tree has been known since [Gower and Ross (1969)](https://doi.org/10.2307/2346439). What I built on top of that is narrower and measurable. The merged reference in `kernels#22` sorts all $n(n-1)/2$ centroid pairs before union-find; sigmoid builds the tree under a strict total order on (distance, lower index, higher index) and runs union-find over its $n-1$ edges. The strict order is what makes the result reproducible on ties: before it, the batch path disagreed with the merged reference on 123 of 200 tie-heavy cases, by up to 4.75 in salience, and random Gaussian inputs, the ones the original parity check used, never exposed it (0 of 40). After it, 0 of 200. The incremental path keeps the tree edges and merges in only the new block's $n$ edges, which the cycle property makes exact under the same strict order. ```python # sigmoid/triton/schedule.py:193-206 @ 71d8a21 for edge in edges: distance, left, right = edge root_left, root_right = find(left), find(right) if root_left == root_right: continue accepted.append(edge) # the smaller component is the one absorbed, and it dies at `distance` if len(members[root_left]) > len(members[root_right]): root_left, root_right = root_right, root_left for block in members[root_left]: salience[block] = distance parent[root_left] = root_right members[root_right].update(members[root_left]) del members[root_left] ``` **Figure 9.** Block salience computed two ways on the same centroids, from all pairs and from the minimum spanning tree, with their maximum difference shown live; the causal block mask built from sink blocks, a local window and the top-k salient blocks, each kept cell labelled with the reason it is kept; and the incremental update when a block is appended, where only the new block's edges are candidates. Tie-heavy presets (integer lattices, duplicated rows) are the families that broke the pre-fix path. The repo's builder timings are shown as a static table. This is the schedule builder only; it does not make the attention faster. Colour key: baseline: the all-pairs reference and its edge count; structure: MST edges kept and blocks selected by topology; parameter: local-window blocks and the order of merges; ink: sink blocks and the centroids; line: blocks the causal mask skips; proof: the two saliences agree. So the claim, bounded by a search I ran in October 2026: I did not find a published system that measures an H₀ barcode of a model's activation window, uses it as the state of a closed-form, action-conditioned operator, reports that operator's measured contraction, and checks every imagined state with a learned restriction-map residual that needs no model call. That is a statement about what I found, not a claim to be first; the nearest work is named above. And, by the numbers above, the combination does not yet predict better than its own null model beyond short horizons. ## Limitations **What failed.** - Withdrawn: "\~30% better rollouts on Lorenz" Killed by: the ablation co-varied PCA rank with topology; against the same-u arm it is about 15% better at k = 1, about 7% better at k = 4, and worse at k = 16. - Withdrawn: "Bit-identical kernel parity" Killed by: only continuous inputs were sampled; 123 of 200 tie-heavy cases disagreed until the strict tie order was added. - Withdrawn: "A useful cheap OOD detector" Killed by: no baselines had been run; mean AUROC 0.846 loses to Mahalanobis 0.899 and PCA reconstruction 0.888. - Withdrawn: "Needs a fixed physical distance" Killed by: a dilating threshold swept over three decades was still read at 0.940 against 0.382 linear; replaced by a scale-relative H₀ plateau. - Withdrawn: Forced contraction (ρ clipped to 0.995) Killed by: one-step NRMSE on Lorenz rose from 0.067 to 0.317. - Withdrawn: Cauchy–Schwarz block pruning for the schedule Killed by: a correct bound that never fires: 0 of 480 blocks pruned. **Prediction.** Apart from short horizons on Lorenz, topology does not improve coordinate prediction anywhere the repo has measured. On the corrected Lorenz table (24-dimensional lift, 2000 steps, 30% held out) sigmoid scores 0.0622, 0.2662 and 1.1266 at $k=1, 4, 16$ against 0.0734, 0.2853 and 0.9889 for the same linear channel without ψ: a delta of −0.0112, −0.0191 and +0.1377. Reading the bench code, the $k=16$ row fails the repo's own topology gate, which requires sigmoid to be below the same-$u$ arm by $10^{-4}$. On distilgpt2 (40 passages, 9186 tokens, 8226 transitions, state dimension 128), sigmoid scores 0.3627 at $k=16$ and the dataset mean scores 0.3616. A layer sweep from the embedding to layer 6 gives deltas against the linear arm between +0.0000 and +0.0025, none negative. On the contact world ψ contributes exactly +0.0000 to coordinate rollout against the same-$u$ arm. Forecasting the topological state itself fails on contact physics: 0.265 at lag 1 against a persistence baseline of 0.250. **Certificates.** The scalar bound is vacuous on distilgpt2 at every setting tried. The directional estimate is tighter but under-bounds by 2.4× to 4.6× when residuals are correlated across steps, which is when one would most want it. **The gate.** It missed uniform random tokens, 0 of 5 fired at a mean score of 0.513, below real prose: it detects structural degeneracy, not unusualness. The effective-rank stalk guards ingestion, not rollout: an imagined state has no window, and the `imagine` path never passes a rank, so the stalk is inert during rollout. The README reports that on scrambled activations the two-term gate fired 0/60 and the stalk 60/60; the test behind those figures uses a stand-in encoder whose ψ and $u$ are both linear in the window mean, not the real encoder. **The representation results, read closely.** The correlation of 1.0000 between β₀ at the true radius and the component count is close to an identity by construction, since the ground truth is defined as Rips H₀ at that radius; the repo says "by construction" itself. The exactness claim against MuJoCo's disjoint-set islands rests on a demo of 12 synthetic two-cluster configurations at one radius, not on MuJoCo. The S²-Rips probe is described as 5-fold logistic regression, but the script in the repo prints a single chronological 70/30 split. The docs describe the linear arm as having the same total state dimension; in the script it is $u$ alone at 16 dimensions, against 40 ψ coordinates under the same settings. And README calls the plateau centre "the interaction radius" while BENCHMARKS says it is not the generator's interaction radius and sits in a band ranked 4th of 19 by width. **Sparse attention.** It does not pay on anything distilgpt2 can run. The attention-op crossover is near 16k positions at 12 heads × 64 and near 8k at 32 × 128, and every winning row is synthesized key/value data at the model's head geometry; only the 1024-position row is a real forward pass. Building the schedule is about 18% of the sparse path at context 960 (44 ms against 199 ms of attention). The inference module's docstring says topology costs 0–20% of decode throughput; the measured end-to-end loss at context 256 is from 102 to 51 tokens per second, about half. $$ q\cdot k=q\cdot c_b+q\cdot(k-c_b)\ \le\ q\cdot c_b+\|q\|\,r_b,\qquad \mathrm{mass}_b\ \le\ n_b\,e^{U_b-L} $$ **Figure 10.** The repo's own negatives in five tabs, each with the control that exposed it. Rollout: the corrected Lorenz and distilgpt2 tables, with the sign flip at k = 16. Layers: the layer-sweep deltas, none below zero. Attention: dense against topology attention-op time over positions, with every row above 1024 marked as synthesized. Pruning bound: the Cauchy–Schwarz centroid-plus-radius bound against the margin a 64-key block would need, which no measured geometry reaches. Radius clause: the five constructed systems that falsified the fixed-physical-distance clause. Colour key: structure: sigmoid and topology arms; measured: controls: same-u and mean arms; the measured gap band; baseline: dense attention, reference bounds and the drift band; withdrawn: losing rows and withdrawn claims. The pruning bound above is correct, and it never fires. To prune a 64-key block at $\varepsilon=10^{-3}$ its ceiling must sit 11.1 nats below the global maximum, while attention logits at scale $1/\sqrt{64}$ span about ±3; the measured minimum of the bound was 2.6e+04 for isotropic keys, 1.88 for tight clusters and 0.47 for very tight ones. It pruned 0 of 480 blocks. It is the same shape of failure as the scalar certificate: true, and vacuous. **Control.** The closed loop works, but the topological cost does not reduce collisions. Over four planar robots, contact radius 0.9 and six episodes each, the β₀ cost gave 11.5 contacts and goal error 1.67; reach-only planning with the topological model gave 10.7 and 1.18; reach-only with the no-topology model gave 9.8 and 1.08. Raising the β₀ weight to 8.0 made it 16.8 contacts. β₀ is exact as an observation but carries a forecast error of 0.113 at horizon 1 and 0.261 at horizon 6, and the planner optimizes differences of about one component against it. **Real time.** The safety check runs at 3.6 µs p50 and 9.8 µs p99 over 20,000 calls, but its maximum was 9458 µs against a 1 ms budget, attributed to CPython garbage collection or scheduling: sub-millisecond at p99, not hard real time, and a real stop belongs outside CPython. Querying `nvidia-smi` costs 32 ms, above the 20 ms contact-rich budget. The real-time module asserts a no-allocation rule while its check expression creates temporary boolean arrays (read from the code, not measured). **Scope.** H₁ is calibration-only, with no measured gain over H₀. Multi-body coupling is verified only on synthetic driver-follower pairs. Nothing is benchmarked against GUDHI, giotto-tda or TDAstats, which were not installed. Automatic configuration was right in 8 of 9 cases, and its recommendation lost to the default in all 3 comparisons on S², both arms above NRMSE 1.0. Checkpoints load with `numpy.load(allow_pickle=True)`, which can execute code from a malicious file; load only checkpoints you produced or trust. **Where the docs and the code disagree** at this commit: - The `state.py` docstring says $u$ is the PCA projection of the window-mean state; `encode` projects the last observation in the window. - The `operator.py` docstring presents the $\rho<1$ projection as the default; the default is `rho_max=None`. - `RolloutCertificate.step_rmse` is documented as measured on held-out data; `CouplingOperator` computes it in-sample and says so. - The `imagine` docstring describes differential decoding; the code decodes $\mathrm{Dec}(u)$ plus the anchor residual decayed by its lag-1 autocorrelation. - The `nbody` module and README describe a tensor-product operator and a gauge projection; the code concatenates body states into one block operator, and `MultiBodyCoupling` clips to 0.995 by default although forced contraction was priced and rejected elsewhere. - The Lorenz sigmoid dimension is 38 in one SIGMOID.md table and 46 in the next table and in BENCHMARKS. - README calls kernel parity "bit-identical, max deviation 1.8e-15"; only the CSR schedules are bit-identical, and salience values deviate by up to 3.6e-15. - Test counts differ across files: 322 in README, 57 in CONTRIBUTING and in the 0.1.0 changelog entry, 21 for `tests/test_sigmoid.py` in SIGMOID.md. Counting test functions at this commit gives 321 across `tests/` and 22 in `tests/test_sigmoid.py`. ## Read more [View the project](https://github.com/teerthsharma/sigmoid) · [Source on GitHub](https://github.com/teerthsharma/sigmoid) The repository is [teerthsharma/sigmoid](https://github.com/teerthsharma/sigmoid), cited as DOI [10.5281/zenodo.21997816](https://doi.org/10.5281/zenodo.21997816). Every line reference in this essay is at commit [`71d8a21`](https://github.com/teerthsharma/sigmoid/tree/71d8a215e85b25c4ce19ec07d141b793f961bc59). The files worth opening first: - [`README.md`](https://github.com/teerthsharma/sigmoid/blob/71d8a215e85b25c4ce19ec07d141b793f961bc59/README.md): the construction, the prior-art table, and section 14 on what I got wrong. - [`SIGMOID.md`](https://github.com/teerthsharma/sigmoid/blob/71d8a215e85b25c4ce19ec07d141b793f961bc59/SIGMOID.md): the study, including the falsification of the radius clause and section 5, "What is established, and what is not". - [`BENCHMARKS.md`](https://github.com/teerthsharma/sigmoid/blob/71d8a215e85b25c4ce19ec07d141b793f961bc59/BENCHMARKS.md): every measured number with its configuration and baseline, including the bad ones. - [`AGENDA.md`](https://github.com/teerthsharma/sigmoid/blob/71d8a215e85b25c4ce19ec07d141b793f961bc59/AGENDA.md): the open problems, the audit of published negative results, and what was deliberately not funded. - [`sigmoid/state.py`](https://github.com/teerthsharma/sigmoid/blob/71d8a215e85b25c4ce19ec07d141b793f961bc59/sigmoid/state.py), [`sigmoid/operator.py`](https://github.com/teerthsharma/sigmoid/blob/71d8a215e85b25c4ce19ec07d141b793f961bc59/sigmoid/operator.py), [`sigmoid/sheaf.py`](https://github.com/teerthsharma/sigmoid/blob/71d8a215e85b25c4ce19ec07d141b793f961bc59/sigmoid/sheaf.py) and [`sigmoid/triton/schedule.py`](https://github.com/teerthsharma/sigmoid/blob/71d8a215e85b25c4ce19ec07d141b793f961bc59/sigmoid/triton/schedule.py): the encoder, the operator and certificates, the gate, and the schedule builder. The open questions the repo states are specific: whether token-to-token is the wrong time axis (generation steps or layers may be the natural one), whether H₁ earns its cost, and an uncertainty-weighted β₀ cost for control in place of more weight. Related essays on this blog: [Aether-Lang](/aether-lang), where the topology-derived sparse-attention kernel merged as `triton-lang/kernels#22` came from; [topological-ml-toolkit](/topological-ml-toolkit), the topology tooling built on Aether-Lang; and, in the same group of essays on attention and learning, [resolvent](/resolvent) and [caustic](/caustic). --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # caustic > A model can hold a fact and still fail to reach it. When it fails, distinct entities collapse onto one answer, and the collapse proves errors. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/caustic - Repository: https://github.com/teerthsharma/caustic - Project site: https://teerth.dev/caustic/ - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "caustic", 2026, https://teerth.blog/caustic - Languages: Python - Question: How many of a model's answers can be proved wrong without knowing a single correct answer? - Headline result: 0.950 certified error floor on 20 country–capital prompts behind 128 tokens of " the", computed from the model's answers alone (control: 128 tokens of coherent prose: floor 0.000; measured error 0.000, against 1.000 behind the repeated token) ## What it is A language model can know a fact and still be unable to reach it. I found this the plain way: I asked `Qwen/Qwen2.5-0.5B` for the capitals of twenty countries, each prompt preceded by a prefix of exactly 128 tokens, and I changed only what those tokens were. With coherent prose in front, the model answered all twenty correctly. With the token `" the"` repeated 128 times, it answered none. The two rows have the same token count and the same entities. The prefix contains none of the answers and is identical for every country, so it carries no information about the task. Only the character of the surrounding text changed. What made this more than a curiosity is the shape of the failure. Under the degenerate prefix the model did not scatter its answers. It gave all twenty countries one answer. Under the failure the model does not become noisy; it becomes constant. Distinct entities collapse onto a single output, and that collapse is something I can see without knowing a single capital. caustic is the package I built around that observation. Its object is the **orbit partition**: take a set of entities that share one relation (countries and their capitals, languages and their countries), ask the model the same question about each, and group the entities by the answer they receive. The groups are the orbits. Nothing in this step needs a correct answer; the partition is computed from the model's own outputs, right or wrong. The result the rest of the project stands on is a counting argument. If the relation is injective, so that distinct entities deserve distinct correct answers, then two entities that share an orbit cannot both be right. An orbit of size $s$ therefore holds at least $s-1$ wrong answers, and over $n$ entities in $m$ orbits at least $n-m$ answers are wrong. That number is a proof, not an estimate, and it needs no answer key. Under the 128 repetitions of `" the"`, twenty countries landed in one orbit, so at least nineteen of twenty answers were provably wrong before anyone looked up a capital. The bound is one-sided, and I want that stated before anything else. It can prove a model wrong. It can never prove a model right. A model that answers every entity with a different wrong answer leaves the partition discrete and the bound at zero. The scope is narrower than "hallucination detection": whether a model is in a regime where retrieval works over a set of entities sharing one relation. It does not score one free-form generation. On the name. The README closes by describing a caustic as a place where distinct preimages merge, and merging is exactly what the partition records. The same line also says the map *folds* there, and I do not lean on that half: the witness map in Theorem 5 below sends distinct points to one image while its Jacobian determinant is positive everywhere, with no fold anywhere. The name is about the merging. **Figure 1.** The same 128 tokens, five ways. For each relation and prefix condition the figure draws the entities as squares, gathers the largest orbit into one block, and computes the certified floor (n − m)/n live from the reported partition. Accuracy and true error need an answer key; the floor does not, and it is never right of the error bar. The right panel shows adjusted Rand index between partitions at different prefix lengths beside the share of answers that changed. Qwen2.5-0.5B, seed 0, one passage per condition. The two relations disagree on shuffled words (0.000 against 1.000), so no row is labelled incoherent for both. Colour key: withdrawn: the largest (collapsed) orbit: the failure state; proof: what the partition alone proves: certified floor, recovery ceiling; structure: partition retained across prefix lengths (ARI); measured: measured: accuracy and true error (need the answer key), share of answers changed. ## What it can do First, it counts errors it cannot see. The floors at fixed prefix length, beside the errors measured with the key: **Measured: 0.250 · 0.000 · 0.950** certified error floor on capital (n = 20): no prefix, 128 tokens of prose, 128 × " the". Control: measured error with the answer key: 0.450 · 0.000 · 1.000; the floor stays below it on every row. n = 20 entities. Source: README.md:534-541 @ 6fed3df · Qwen2.5-0.5B, float32, seed 0, RTX 4060 Laptop. On `language` (twelve entities) the same three conditions give floors of 0.333, 0.000 and 0.833 against measured errors of 0.500, 0.250 and 1.000. The prose row on `language` is the bound behaving as a bound should: the floor went to zero while three answers were still wrong. On `capital` the choice of prefix moves the provable floor across a 95-point range at identical token count, and the move can go the wrong way: going from no prefix to the degenerate one adds 0.700 to the floor. **A sharper count, still without a key.** Theorem 1 sees only collisions. If I also hand the certificate the correct answers *as a set*, without the pairing (that Paris and Tokyo are capitals, not which country owns which), it can also count answers that are nobody's correct answer. I call that Theorem 1\* and state it in the next section. Across three models it matched the true error count exactly on every evaluable row: **Measured: 16 / 16** rows where Theorem 1\* equals the true error count (distilgpt2, SmolLM2-135M, Qwen2.5-0.5B; capital, language, currency; two prefixes). Control: Theorem 1 on the same rows: strictly below the true count in 16 of 16; 60 bound checks, 0 violations, 4 relation–model pairs skipped by the injectivity check. n = 16 evaluable rows. Source: RESULTS.md:382-403 @ 6fed3df · true error computed from gold after the bound. The row I care about is `currency` on Qwen with no prefix: twelve entities, twelve distinct answers, so Theorem 1 certifies zero, while six answers are wrong and Theorem 1\* certifies exactly six. The exactness needs a control, and the repository supplies one against itself: eight of the sixteen rows are the `" the"`×128 condition, where the model emits one token that is nobody's answer, so $m^* = 0$ and any sound bound is exact by construction. On the other eight rows, RESULTS.md says Theorem 1 is "loose by exactly one". Its own table says otherwise. Reading the rows in table order, Theorem 1 trails the true count by 1, 1, 3, 2, 1, 4, 2 and 6 answers. I report the table. **Repair, measured rather than asserted.** `repair_by_context` partitions the relation twice, without and with a prefix, and reports the partition on both sides. It calls a run `REPAIRED` only for a collapsed-to-discrete transition and `WORSENED` whenever the prefix merges entities that were separate. The accuracy column appears only when a caller passes `gold`, and the verdict never consults it. The prefix I ship, `NEUTRAL_PREFIX`, is 128 tokens on mechanical calculators, ocean currents, language and photosynthesis. Its effect on `capital` is reported two ways inside the same repository, and I am not going to pick one: README §10.1 says it is worth 0.550 → 1.000, while the docstring in `repair.py` and the quickstart notebook give 0.750 for the exported constant, and the README's own scope section agrees with 0.750 and attributes the 1.000 to a longer tiled passage. A cleaner separation came from constraining the decode. Restricting the argmax to the set of correct answers takes accuracy from 0.550 to 1.000 on `capital`, 0.625 to 1.000 on `language` and 0.500 to 1.000 on `currency` with no prefix; under `" the"`×128 the same restriction gives 0.050, 0.062 and 0.000. Without the degenerate prefix the errors were the model leaving the answer space. Under the degenerate prefix the fact itself is unreachable. Constraining changes the task toward multiple choice, so those accuracies are not comparable to free decoding. **What the model holds when it is wrong.** On 100 wrong items, the correct answer had median rank 3 of 151,936 vocabulary entries, sat within the top 10 in 88 and within the top 1000 in all 100, and the entity was still linearly recoverable from layer 22 at 0.9624 (chance 0.0312), higher than on correct items (0.8981). On the 15 wrong items of a 32-item set, the chosen token led the correct one by a mean logit gap of 0.8338, which is 0.2526 of the logit standard deviation over the vocabulary. The model is not confidently wrong. It is wrong by a quarter of a standard deviation, which is why a weak intervention can move it so far. **Noise as an intervention, scored by the floor.** A near-tie is the condition under which noise can help a readout. I added Gaussian noise to the input embeddings, scaled by their own standard deviation, and took a majority vote over 16 draws per level. $$ \tilde{x} = x + \sigma \cdot \mathrm{std}(x) \cdot \varepsilon, \qquad \varepsilon \sim \mathcal{N}(0, I) $$ **Measured: 0.500 → 0.583 → 0.833** accuracy on language (n = 12) at σ = 0.00, 0.40, 0.80; distinct answers 8 → 9 → 11; certified floor 0.333 → 0.250 → 0.083. Control: σ = 0 baseline; gain +0.333 at σ = 0.80. Under identical noise capital declines monotonically, 0.550 → 0.200. n = 12 entities, 16 votes per level. Source: README.md:663-669, 717-720 @ 6fed3df · Qwen2.5-0.5B, seed 0; σ grid of caustic/experiments/stochastic_resonance.py:50. I quote only the σ values the committed script actually runs. Its grid is `(0.0, 0.01, 0.02, 0.05, 0.10, 0.20, 0.40, 0.80)` and stops at 0.80. The README table continues to 1.20, 1.60 and 3.00, with accuracy at 0.667 at 1.20 and zero at 1.60 and 3.00, but that sweep is not the one in the repository, and the committed script, given a peak at its last grid point, prints that accuracy is still rising and that the sweep should be extended. So on the three README rows whose σ the committed grid contains, accuracy rises and the floor falls together up to the edge of the grid, and the fall on the far side is not established. On `capital` noise only hurts. There is no good $\sigma$ to recommend, only a procedure: sweep it and let the label-free floor choose. **System prompts and ensembles.** A system prompt is a prefix, so it can move the partition too. The README counts six prompt conditions on `capital`, the empty prompt among them as the reference. None of the five actual prompts left the partition intact: adjusted Rand index against the no-prompt partition was at most 0.2721, and exactly 0.0000 for two of them. The prompt asking the model to say so rather than guess when uncertain cut accuracy from 0.550 to 0.300 and merged orbits from 15 to 9, raising the certified floor from 0.250 to 0.550. A JSON-format instruction reached 0.950. Averaging logits across eight neutral prefixes scored 0.100 on `capital`, below the 0.550 of one pass, while majority vote scored 0.600. ## How it was made The construction has one object, a handful of counting theorems over it, and four results from other branches of mathematics that explain why the counting is the right thing to measure. The theorems are proved on paper in the module docstring of `caustic/theorems.py`, each with an executable witness; the repository has no Lean. > **Definition: Orbit partition.** > > Let $E$ be a finite set of entities, $|E| = n$, and let $f : E \to A$ be the model's answer map. The orbit partition $P(f)$ is the partition of $E$ by the value of $f$, whose blocks (orbits) are the classes of $e_1 \sim e_2 \iff f(e_1) = f(e_2)$. Write $m = |P(f)|$ and $k_e$ for the size of the orbit containing $e$. It is $H_0$ of the relation "these two entities receive the same answer". Source: README.md:262-268, 405-407 @ 6fed3df. $$ e_1 \sim e_2 \iff f(e_1) = f(e_2) $$ **Figure 2.** Eight entities, twelve answer slots: the eight capitals and four tokens that are nobody's capital. Moving an entity's cube onto an answer stacks it there, so stack height is orbit size, and the inset draws the H0 graph whose components are the orbits. The partition is computed whether the answers are right or wrong; the key toggle only adds check marks. France, Japan, Peru, Kenya and Norway are the README quickstart's entities; Italy, Egypt and Chile are from the repository's coupling_gap table. The toy is not a model, and height is a count, not a network quantity. Colour key: withdrawn: an orbit of size greater than one (collapse); structure: an answer that lies in the set of correct answers G (solid ring), and the check mark the key toggle adds; ink-3: dashed ring: an answer that is nobody’s correct answer; ink: a cube alone on its answer (orbit of size one). The code that computes it is a dictionary. This is the end of `orbit_partition` at the pinned commit: ```python # caustic/regime.py:293-304 @ 6fed3df orbits: dict[int, list[str]] = {} for e, a in zip(spec.entities, answers): orbits.setdefault(a, []).append(e) counts = [len(v) for v in orbits.values()] return OrbitReport( entities=tuple(spec.entities), answers=tuple(answers), injective=spec.injective, n_distinct=len(orbits), largest_orbit=max(counts), orbits=orbits, ) ``` The answers are typically argmax token ids for the first template. A batched path re-runs the shortest prompt alone and raises on a mismatch, which catches right padding. > **Theorem: Theorem 1 — Orbit Error Bound (proved).** > > For an injective relation $R : E \to A$, the number of entities on which $f$ errs is at least $n - m$. Tight exactly when every orbit contains one correct answer. > > *Proof (as in the repository).* If $e_1 \neq e_2$ lie in one orbit then $f(e_1) = f(e_2)$ while $R(e_1) \neq R(e_2)$, so $f$ is wrong on at least one of them. An orbit of size $s$ contributes at least $s - 1$ errors; summing over the $m$ orbits gives $\sum_i (s_i - 1) = n - m$. Source: caustic/theorems.py:17-25 @ 6fed3df. $$ \mathrm{err}(f) \geq n - m $$ **Figure 3.** Every proved bound on one assignment, against the true error count. The error bar is split into what Theorem 1 proves from collisions, the extra that Theorem 1\* proves from the answer set, and the wrong answers no bound can see. Below it, the precision floors of Theorems 6 and 6\* and the recall floor of Theorem 8 against realised values; a stress test draws 2000 random instances and counts violations. The shuffle preset shows the case the certificate cannot see: nothing collides, nothing is inadmissible, the recall floor is zero. Colour key: proof: proved bound (darker: Theorem 1 and the Theorem 6 and 8 floors; lighter: Theorem 1\*, the sharper bound); measured: true error and realised precision and recall (need the key); baseline: slack of the weaker Theorem 1 bound (left histogram; the right one, slack of Theorem 1\*, is proof). In code the certificate is one subtraction, guarded by the precondition: ```python # caustic/regime.py:145-147 @ 6fed3df if not self.injective: return 0 return len(self.entities) - self.n_distinct ``` Theorem 1 cannot see an entity that answers alone with something that is nobody's correct answer. The sharper bound uses the answer set $G = R(E)$, which is a set and not a key. > **Theorem: Theorem 1\* — Admissible Orbit Error Bound (proved).** > > Let $G = R(E)$ be the correct answers as a set and $m^* = |f(E) \cap G|$. For injective $R$, at least $n - m^*$ answers are wrong, and $m^* \le m$, so this is never weaker than Theorem 1. > > *Proof.* Let $C = \{e : f(e) = R(e)\}$. On $C$, $f = R$ is injective, so $|C| = |f(C)|$, and $f(C) \subseteq f(E) \cap G$, so $|C| \le m^*$. Source: caustic/theorems.py:33-42 @ 6fed3df. $$ m^* = |f(E) \cap G|, \qquad \mathrm{err}(f) \geq n - m^* $$ **Measured: 0 → 6** certified errors on currency, Qwen2.5-0.5B, no prefix: Theorem 1, then Theorem 1\*. Control: true error count from the gold answers: 6. n = 12 entities. Source: RESULTS.md:393-405 @ 6fed3df · template 0, seed 0. `admissible_distinct` is where $m^*$ is computed, and passing `n_entities` lets the certificate check its own precondition, because $|G| \lt n$ at the compared encoding is exactly a failure of injectivity: ```python # caustic/regime.py:452-459 @ 6fed3df gold = set(gold_keys) if n_entities is not None and len(gold) < n_entities: raise ValueError( f"gold_keys holds {len(gold)} distinct values for {n_entities} " "entities, so the relation is not injective at this encoding and " "Theorem 1* does not apply; see regime.verify_injective" ) return len(set(answers) & gold) ``` A count is not actionable; a caller has to know which entities to withhold. Theorem 6 turns the count into a set with a precision floor proved before it is evaluated. Let $S = \{e : k_e \gt 1\}$ be the entities in non-singleton orbits and $b$ the number of such orbits. > **Theorem: Theorems 6 and 6\* — Certified-set precision (proved).** > > For injective $R$, $\mathrm{precision}(S) \ge (n - m)/|S| = 1 - b/|S|$. With $b_{adm}$ the number of non-singleton orbits whose shared answer lies in $G$, $\mathrm{precision}(S) \ge (|S| - b_{adm})/|S|$, and $b_{adm} \le b$. Because $|S| \ge 2b$, the Theorem 6 floor is never below one half on a non-empty set. Sources: caustic/theorems.py:194-205, 226-237 @ 6fed3df. $$ \mathrm{precision}(S) = \frac{|S \cap \mathrm{wrong}|}{|S|} \geq \frac{n - m}{|S|} = 1 - \frac{b}{|S|}, \qquad \mathrm{precision}(S) \geq \frac{|S| - b_{adm}}{|S|} $$ **Measured: 0.950 → 1.000** precision floor on capital under 128 × " the": Theorem 6, then Theorem 6\* (b_adm = 0). Control: realised precision of the withheld set: 1.000. n = 20 entities. Source: caustic/theorems.py:492-497 @ 6fed3df · Qwen2.5-0.5B, seed 0. Across 80 conditions (two models, four relations, five templates, two contexts) the docstring reports zero violations, and realised precision 1.000 in 56 of the 60 evaluable conditions against a mean proved floor of 0.756. For a long time I believed recall could not be bounded at all, and the README still says so in two places. That was wrong. The withheld set $S^*$ is $S$ extended by every entity whose answer lies outside $G$; Theorem 1\* puts all of its certified errors inside $S^*$. > **Theorem: Theorem 8 — Recall floor (proved).** > > Whenever $f$ errs at all, $\mathrm{recall}(S^*) \ge (n - m^*)/n$. *Proof.* Theorem 1\* places all $n - m^*$ certified errors inside $S^*$, so $|S^* \cap \mathrm{wrong}| \ge n - m^*$, and $|\mathrm{wrong}| \le n$. The floor is zero exactly when $m^* = n$. Source: caustic/theorems.py:295-302, 314-317 @ 6fed3df. $$ \mathrm{recall}(S^*) = \frac{|S^* \cap \mathrm{wrong}|}{|\mathrm{wrong}|} \geq \frac{n - m^*}{n} $$ **Measured: 2,048,574** configurations checked exhaustively (every answer map over n + 2 symbols, every injective truth, n = 2..5, at least one error). Control: 0 violations; tightest margin exactly 0 (the bound is attained); floor strictly positive in 99.3%. Source: RESULTS.md:425-431; tests/test_recall_no_go.py:145-183 @ 6fed3df. Its zero is Theorem 7's witness: a model that shuffles the correct answers among the entities. The partition is discrete, nothing is inadmissible, nothing is withheld, and recall is 1 under one consistent truth and 0 under another. So no *constant* positive recall floor exists, and Theorem 8's floor is tight rather than absent. **What a single answer destroys.** If a block of $k$ entities shares one answer, nothing downstream can tell them apart. > **Theorem: Theorem 2 — Pooling Recovery Bound (proved).** > > If $f$ maps a block of $k$ entities to one answer, then for any downstream $h$, the probability that $h(f(e))$ recovers $e$ under a uniform prior on the block is at most $1/k$. *Proof.* $h \circ f$ is constant on the block, so it agrees with the identity on at most one of its $k$ entities. Source: caustic/theorems.py:59-67 @ 6fed3df. $$ \Pr[\, h(f(e)) = e \,] \leq \frac{1}{k} $$ **Figure 4.** Part A enumerates every answer map and every decoder on k ≤ 4 entities (65,536 pairs at k = 4) and shows that the best decoder recovers exactly one entity per orbit, never more than 1/s inside an orbit of size s. Part B gives the decoder every paraphrase: the join of the per-template partitions is at least as fine as any one of them, so the Theorem 2\* ceiling m_join / n is never lower than the single-template one; the eight-entity answer tables in part B are illustrative shapes of the measured rows. Measured Qwen rows: capital without prefix, m_join 20 and ceiling 1.000 (single template 0.250); under 128 × “ the”, m_join 1 and ceiling 0.050. Colour key: proof: proved ceiling; baseline: single-answer ceiling, the weaker reading; measured: enumerated decoders and measured rows. Theorem 2\* is the same argument on the full answer tuple across $T$ templates: a receiver seeing every paraphrase recovers at most $m_{join}/n$ entities. Measured on Qwen over five paraphrases, the join was discrete for every injective relation under coherent context and coarse under the degenerate prefix, with ceilings of 0.050 to 0.167. That split separates a repairable regime from an unrepairable one. **Why collapse happens.** Two results connect the partition to the model's geometry. Neither is a detector. > **Theorem: Theorem 3 — Zero coupling implies pooling (proved).** > > Let $z_c(h)$ be the logit of token $c$ as a function of the entity representation $h$, continuously differentiable on a domain containing a path $\gamma$ from $h_1$ to $h_2$. If the directional derivative of $z_c$ along $\gamma$ vanishes identically, $z_c(h_1) = z_c(h_2)$; if this holds for every candidate $c$, the two entities share an orbit. Source: caustic/theorems.py:123-136 @ 6fed3df. $$ z_c(h_2) - z_c(h_1) = \int_{\gamma} \nabla z_c \cdot d\ell = 0 $$ **Figure 5.** A two-dimensional stand-in for the entity representation, with three candidate tokens whose logit surfaces the reader shapes. The figure evaluates the path integral the way caustic/theorems.py does (midpoint rule, 2048 steps) against the closed form, draws the directional derivative along the path, and reports whether h₁ and h₂ receive the same answer. Removing the coupling along the path for every token gives equal logits and a shared answer. The surfaces are illustrations of the identity, not the model's logits. Colour key: structure: logit surface, logit bars, area under g(t) (the integral); measured: path, markers h₁ and h₂, g(t) and its arrows along the path, argmax mark, readouts; proof: equality badge: integral equals the logit change. The only measured shadow of this theorem is a finite-difference ratio on wrong items: the token the model chose coupled to the entity at 0.93 times its coupling to control tokens, against 1.33 for the correct token. The partition observes the consequence whether or not the coupling is measurable, which is why the partition is what I measure. > **Theorem: Theorem 4 — Dissipative pooling (proved).** > > Let $T$ be differentiable with characteristic exponents $\lambda_1 \ge \dots \ge \lambda_D$ and $S = \sum_i \lambda_i \lt 0$. For any bounded $A$ of positive Lebesgue measure, the volume of $T^n(A)$ tends to zero at rate $e^{nS}$. *Proof.* $\mathrm{vol}(T^n(A)) = \int_A |\det D(T^n)|$ and $(1/n)\log|\det D(T^n)| \to S$. Source: caustic/theorems.py:144-153 @ 6fed3df. $$ \mathrm{vol}(T^{n}(A)) \sim e^{nS} \longrightarrow 0 $$ **Figure 6.** A dissipative map iterated on twenty points. The exponents are computed by the repository's QR algorithm (Benettin re-orthonormalisation), the tangent ellipse of the starting disc is drawn at every fifth step, and its area is compared with e^(nS). A slice at the current step counts how many decision cells the twenty points occupy and computes (20 − m)/20. The Hénon map and the linear contraction are toys that satisfy the theorem exactly; they are not models of any network. Colour key: structure: tangent ellipses, cell occupancy, and the nS reference dots; measured: nonlinear trajectories, the 20 current points, and the computed exponents and area line; proof: floor from occupied cells, and the area identity. **Measured: S = −226.74** sum of finite-time characteristic exponents of the token-position Jacobian product, block 3, 46 steps; λ₁ = +0.1653, 139 of 768 directions expanding. Control: shuffled-token control: S = −170.37, λ₁ = +0.1852, 151 of 768 expanding. Source: README.md:1010-1013 @ 6fed3df · distilgpt2 (D = 768), not the Qwen model of every other result here. That measurement is on `distilgpt2`, a different network from the one whose collapse I report, and the repository forbids combining the two. The exponents are finite-time quantities over 46 steps, not Oseledets limits, and the module says so. They supply the hypothesis $S \lt 0$ for a network; they do not separate correct from wrong answers. **Why I stopped looking at the Jacobian.** My first route to detecting this failure was spectral: summaries of the layer-to-layer Jacobian such as its largest singular value, log volume and tail exponent. That arm sat at chance. Theorem 5 is the reason it had to. > **Theorem: Theorem 5 — No local criterion detects pooling (proved).** > > There is a smooth map whose Jacobian is nonsingular at every point and which is not injective. Consequently no function of the local Jacobian alone (determinant, smallest singular value, condition number, spectral decay or any other pointwise invariant) can decide injectivity. *Witness:* $\det DF = e^{2x} \gt 0$ everywhere, yet $F(x,y) = F(x, y + 2\pi)$, and the Jacobians at the two preimages differ by a rotation, so they share every spectral invariant. Source: caustic/theorems.py:165-176 @ 6fed3df. $$ F(x, y) = (e^{x}\cos y, \; e^{x}\sin y) $$ **Figure 7.** The covering surface of the Theorem 5 witness: each sheet is one turn of y, lifted by height. Two points on different sheets (adjacent by default, winding j sets 1 to 3 turns apart) project to one image point, and the table compares the Jacobian at both: determinant, singular values, condition number, all equal to rounding. Nothing singular marks the collision; there is no fold. A detector that watches a pointwise Jacobian statistic reads the same numbers at both points. The shipped detector compares two entities, which is the global step this theorem says is needed. Colour key: parameter: amber ramp, light to dark: sheet index k, one turn of y per sheet; measured: the two preimages and their shared image; baseline: the pointwise-Jacobian reading this theorem rules out; proof: invariants identical at both points. The escape is global: compare two entities rather than examining one point. On cost, the full $768 \times 768$ Jacobian of one `distilgpt2` block took 53.694 ms against 0.588 ms for the block's forward pass, and the Jacobian route reached 0.61. The shipped detector is five forward passes per entity and no derivative of anything. **Using the floor as an objective.** Because the floor needs no key, I can rank interventions by it at inference time. `select_prefix` always enters the empty prefix under the name `none`, partitions under each candidate, and declines unless a candidate is strictly better: ```python # caustic/governor.py:218-225 @ 6fed3df # Decline unless a candidate strictly beats doing nothing. best = min( (nm for nm in pool if nm != "none"), key=lambda nm: (scores[nm], list(pool).index(nm)), default=None, ) if best is None or scores[best] >= baseline: best = "none" ``` It refuses a non-injective relation, and returns at once when the baseline is already discrete, since a zero floor cannot be beaten strictly. `guard` calls it, withholds the entities in shared orbits (and, given the answer set, those answering outside it), and reports the precision floor of what it withheld. The README credits constructions from my earlier repositories: the seeded Johnson–Lindenstrauss frame from Epsilon, the resonance framing from epsilon-cli, the noise sweep from EPSILON-PHASE, the competition between candidates from laamba-silence. What is new is what they are pointed at: a scorer that needs no ground truth. ## What's new in it The usual way to detect a wrong answer without ground truth is self-consistency: ask again, under sampling or paraphrase, and distrust answers that disagree with themselves. caustic includes that half on purpose as the baseline. It calls it invariance: paraphrase the prompt and the answer must not change. What it adds is the other half of the symmetry a fact carries, equivariance: swap the entity and the answer must change. A model outside its retrieval regime is invariant where it should be equivariant, giving the same answer whichever country is named, and self-consistency scores that state as healthy. The per-entity collision score is the fraction of *other* entities that receive this entity's answer. **Measured: 0.995 \[0.97, 1.00]** AUROC of collision (equivariance) for flagging wrong answers on capital, an injective relation whose errors are collapse. Control: invariance (self-consistency) on the same items: 0.859. On harder relations whose errors disperse, pooled collision AUROC is 0.7083 \[0.3809, 1.0000] (short context) and 0.6687 \[0.3819, 0.9167] (full context): both intervals include 0.5. Interval: 95% percentile bootstrap. n = 20 entities on capital; 32 pooled with language; minority classes 1, 3, 4, 8. Source: RESULTS.md:85-89, 336-350 @ 6fed3df · Qwen2.5-0.5B, seed 0. The 0.995 is the method's score against the cause it was built for, and I read it that way, not as a hallucination-detection figure. It detects collapse, not error. On `language` collision scored 0.950 against invariance 0.942. Where each entity is wrong in its own way, nothing collides, nothing fires, and the pooled AUROC drops to 0.67–0.71 with intervals that span chance. At these sizes (n = 12–20) a 95% bootstrap interval has power 0.00–0.38 to separate a true 0.70 from 0.50, so what the design can say is that it cannot resolve whether the detector degrades. The second difference is where I stopped expecting a score at all. An AUROC needs labels to compute and forces a proved quantity and an unproved one into a single number. The certificate is a deterministic inequality over a finite set: no sampling distribution, no minority class, nothing to be underpowered about. Its precision is proved before it is evaluated, its recall has a floor that is attained, and its silence is a theorem rather than a weakness of the argument. That is why the headline of this project is the bound and not the AUROC. The third difference is against spectral and Jacobian detectors: a determinant, a smallest singular value or a condition number is a quantity Theorem 5 proves cannot see collapse, and the repository's own Jacobian arm is the measured instance. The fourth difference is against tuning an intervention on held-out accuracy: the floor is a number a deployed system can compute, so the prefix, the noise level and the decision not to intervene are chosen without labels. **Figure 8.** Tab A re-runs the Theorem 8 check in the browser: every answer map over n + 2 symbols and every injective truth, n = 2..5, counting violations and the margin between realised recall and the floor (the full run is 2,048,574 configurations); the same enumeration also checks Theorems 1 and 1\*, an addition of this figure. The Theorem 7 witness shows one observation under two truths, recall undefined and recall zero, with the floor at zero. Tab B draws the sixteen model rows with their true error counts; the eight 128 × “ the” rows are hatched, because there any sound bound is exact by construction. Colour key: proof: proved bound (outlined: Theorem 1; filled: Theorem 1\*), violation counter, margin histogram; measured: true error, computed from gold after the bound; baseline: slate hatch: rows with m\* = 0, where any sound bound is exact by construction. ## What no one else built I compared caustic against the closest work I could find. Each line names the work and the concrete difference. **Self-consistency.** [Wang et al., Self-Consistency Improves Chain of Thought Reasoning in Language Models (arXiv:2203.11171)](https://arxiv.org/abs/2203.11171) samples several reasoning paths and takes the answer they agree on most. It is a decoding rule for one question. caustic's invariance half is close to this and is shipped as the baseline. The difference is that caustic compares answers *across entities* of one relation, not across samples of one question, and the agreement it penalises is between different entities. **SelfCheckGPT.** [Manakul, Liusie and Gales, SelfCheckGPT (arXiv:2303.08896)](https://arxiv.org/abs/2303.08896) samples more responses from a black-box model and flags sentences the samples contradict, with no external database. It is zero-resource like caustic, and it is evaluated by AUC-PR against human annotations. It returns a score per sentence; caustic returns a count of errors that is provably a lower bound, plus a withheld set with a proved precision floor, from deterministic answers rather than samples. **Semantic entropy.** [Farquhar, Kossen, Kuhn and Gal, Detecting hallucinations in large language models using semantic entropy (Nature 630, 2024)](https://www.nature.com/articles/s41586-024-07421-0) clusters sampled answers by meaning and computes entropy over the clusters; high entropy flags confabulation. The authors note that the method does not directly address cases where a model is confidently wrong. Semantic entropy measures uncertainty within one question. caustic looks across questions: a model that gives twenty countries one capital shows maximal collision whatever its uncertainty on each question, and the certificate counts the errors that implies. **Metamorphic testing.** [Yang, Al Mamun, Zhang and Uddin, Hallucination Detection in Large Language Models with Metamorphic Relations (arXiv:2502.15844)](https://arxiv.org/abs/2502.15844) mutates prompts and treats a violated metamorphic relation as a sign of hallucination, without external resources, reporting F1 against SelfCheckGPT. Entity swapping under an injective relation is a metamorphic relation in that sense. What the paper does not give, and caustic does, is a theorem that turns the violations into a lower bound on the number of wrong answers, and a precision and recall floor for the set it flags. **Probing for truthfulness.** [Azaria and Mitchell, The Internal State of an LLM Knows When It's Lying (arXiv:2304.13734)](https://arxiv.org/abs/2304.13734) trains a classifier on hidden activations with labelled true and false statements. [Burns, Ye, Klein and Steinhardt, Discovering Latent Knowledge in Language Models Without Supervision (arXiv:2212.03827)](https://arxiv.org/abs/2212.03827) finds a direction in activation space without labels by requiring a statement and its negation to receive opposite truth values. Both read internal states; the first needs labels to train. caustic reads only the model's outputs, and its guarantee does not depend on how well any probe generalises. caustic's own probe result is compatible with the premise of that line of work, that internal states hold more than the output shows: on wrong items the entity was still linearly recoverable at 0.9624, and the correct token ranked a median 3rd. **Guarantees with calibration.** [Mohri and Hashimoto, Language Models with Conformal Factuality Guarantees (arXiv:2402.10978)](https://arxiv.org/abs/2402.10978) gives high-probability correctness guarantees by conformal prediction, backing off to less specific outputs, using a small set of human-annotated samples. Its guarantee is probabilistic and needs that calibration set. caustic's floor is deterministic, needs no annotated samples, and is one-sided: it proves errors, never correctness. **Theory of why hallucinations happen.** [Kalai, Nachum, Vempala and Zhang, Why Language Models Hallucinate (arXiv:2509.04664)](https://arxiv.org/abs/2509.04664) analyses hallucination as a consequence of binary-classification error under the statistics of pretraining and of evaluations that reward guessing. That is an account of causes in training. caustic's bounds are evaluated at inference, on one model's answers over one relation. **The counting itself.** The core of Theorem 1 is the pigeonhole principle; the classical form is that two objects with identical answers cannot both be identified ([Shor, 18.310 lecture notes on the pigeonhole principle, MIT](https://math.mit.edu/~shor/18.310/pigeonholenotes.pdf)). I claim nothing about the counting. What I built is its use as an inference-time instrument for a language model, and what survives the comparison above is this combination: the orbit partition of an injective relation as the observable; Theorems 1 and 1\* as a certified lower bound on the number of wrong answers, computed with no key and at most the answer set; Theorems 6, 6\* and 8 as proved precision and recall floors for the withheld set, with Theorem 7 marking exactly where the recall floor is zero; Theorem 5 as the reason a pointwise Jacobian cannot replace it; and the floor used as the objective that selects or declines an intervention. In the work I compared against I did not find a deterministic lower bound on error count of this kind; I do not claim more than that. **Figure 9.** Three families of intervention scored by one label-free number. For the noise level, the prefix and the system prompt, the figure plots the certified floor (no key) beside accuracy (needs a key), adds the number of distinct answers on the noise panel and the adjusted Rand index against no prompt on the system-prompt panel, then runs the select_prefix rule: empty candidate always entered, ties to the earlier entry, decline unless strictly better. Rows whose floor is above doing nothing are marked. The σ panel uses the README rows; only σ = 0.00, 0.40 and 0.80 are on the grid of the committed script. A two-logit sketch illustrates the 0.2526-sd near-tie and is not the model; a cost panel gives sequential and batched timings for 20 prompts in neutral bars. Colour key: proof: certified floor, computed with no answer key; measured: accuracy, which needs the key (also the shaded swap tail in the near-tie sketch); structure: distinct answers m and partition ARI; withdrawn: an intervention worse than doing nothing. ## Limitations The certificate is one-sided, the detector detects collapse, and almost every measurement rests on two small relations and one model under half a billion parameters. Here is what failed and what is not covered, as the repository states it. **What failed.** - Withdrawn: Collision AUROC 0.995 as a hallucination-detection figure Killed by: harder relations: pooled 0.7083 \[0.3809, 1.0000] and 0.6687 \[0.3819, 0.9167], both intervals include 0.5 (RESULTS.md:336-362). - Withdrawn: Theorem 7 as first stated: no recall floor exists Killed by: Theorem 8, recall(S\*) ≥ (n − m\*)/n, checked over 2,048,574 configurations with 0 violations (caustic/theorems.py:256-257, 295-317). - Withdrawn: Bootstrap AUROC intervals from the first auroc_ci Killed by: single-class resamples were scored 0.5, pinning the percentile bounds; the current code discards them (caustic/experiments/ci_recompute.py:1-12). - Withdrawn: Jacobian spectral summaries (sigma_max, log_volume, tail_alpha) as a detector Killed by: sat at chance, 0.61 at 53.7 ms per position; Theorem 5 says no pointwise invariant can work (caustic/regime.py:3-5; triangulate.py:3-4). - Withdrawn: An earlier geometric gate Killed by: 0.846 mean AUROC against Mahalanobis 0.899 and PCA 0.888, both cheaper (caustic/detect.py:9-12). - Withdrawn: K1: J-space summary statistics as a hallucination detector Killed by: shuffled-token control: signs flip layer to layer for three of four statistics (CANDIDATES.md:475-501). - Withdrawn: K2: Oseledets/persistence bridge on a transformer cocycle Killed by: tolerance sweep: entropy/logD = 0.9986 at the finest tolerance, bar count slides 763 to 1 with no plateau (CANDIDATES.md:263-283). - Withdrawn: ARI-filtered template averaging as a fix for the averaged score Killed by: ARI–AUROC correlation −0.111 over 20 pairs; filtering to ARI ≥ 0.5 made three relations of four worse (template_agreement.py:1-14). - Withdrawn: Mean-logit ensembling over eight prefixes Killed by: 0.100 on capital, below the 0.550 single pass (README.md:803-816). - Withdrawn: NEUTRAL_PREFIX as a passage with no chemistry Killed by: it names chemical energy, carbon dioxide, oxygen and a reaction; a stripped passage still repairs element_symbol at 0.875 (README.md:72-76). - Withdrawn: A universally good noise level Killed by: capital declines monotonically under the same noise, 0.550 → 0.200 (README.md:717-725). **Figure 10.** Where the detector works and where it does not. Panel A plots every reported collision AUROC with its interval, and the invariance baseline as open rings, against the chance line: capital and language, where errors are collapse, then the harder relations, where intervals cross 0.5, and the many-to-one relation, where the sign inverts. Panel B is a live twelve-entity toy: pick collapse, dispersed or shuffle answers and read AUROC with a bootstrap interval (single-class resamples dropped) beside the certificate. Panel C shows numeric answers under a tokenizer that splits off the space: twenty numbers share one first token, and Theorem 1 would certify nineteen wrong on twenty hypothetical correct answers. As failures move from collapse to confusion, the certificate goes silent. Colour key: measured: measured or live AUROC with its interval (open ring: invariance baseline), true error; proof: certificate; withdrawn: coral: the invariance baseline (live), points below chance, the part of an interval on the chance side, answers sharing a token (the tokenizer failure). **The precondition is token-level.** Theorem 1 needs the relation to be injective in the encoding the answer function compares, which for a top-1 token detector means distinct gold answers must have distinct first tokens. `small_capital` violates it: Asmara and Asuncion share a token, as do Lusaka and Ljubljana. No published number is affected (0 of 14 certified errors on that relation are collision artifacts), but the precondition held by luck rather than construction. Numeric answers are worse: under the Qwen tokenizer `" 20"` becomes `[220, 17, 15]`, every number shares first token 220, and Theorem 1 would certify nineteen wrong answers on a model that answered all twenty correctly. `verify_injective` raises on exactly this. On a genuinely many-to-one relation (`continent`) the collision signal inverts to 0.273, and `select_prefix` raises. **Injectivity is necessary and not sufficient.** `currency` is injective at the token level and its averaged collision AUROC is 0.3056, below chance and below each of its components, because the score averaged five templates while the label came from one. `collision_scored` reports the scored template alone. One accuracy figure, `element_symbol` at 1.000, is graded on first letters for four of its sixteen golds and is not established. **Scope of the evidence.** The retrieval, partition, detection, noise, prompt and ensemble results are on `Qwen2.5-0.5B`; the dynamics and cost are on `distilgpt2`; the two must not be combined, and reconciling them is open work. There are two injective relations of 12 and 20 entities, one distractor passage per condition, one seed. All three models in the cross-model table are base models under 0.5B parameters, where the dominant failure is collapse. RESULTS.md calls this the largest limit on the page: as failures shift from collapse to confusion (a plausible answer belonging to the wrong entity), orbits go discrete, the certified set empties and the certificate goes silent, and nothing here tested a model where that could be observed. Whether coherence-gated retrieval is a general property of language models is not established. Answers are compared by top-1 token, so a correct answer phrased differently counts as disagreement. Injectivity was diagnosed after the failure on a many-to-one relation, not predicted. The bounds constrain a partition and say nothing about how often such partitions arise in deployment. **A non-reproduction.** In the quickstart notebook, at 135M parameters (SmolLM2-135M), the headline contrast does not reproduce. Both prefixes collapse the partition, and coherent prose collapses it harder: every entity lands on `" the"`. The same notebook's Qwen prose row reads 0.950 against the 1.000 in RESULTS.md, because the two runs use different 128 tokens of prose. **Experiments with no result.** `guard_at_inference` (the guard against one-forward-pass baselines at matched coverage), `dky_predicts_pruning`, `wrongness_auroc`, `fold_collision` and `attention_to_entity` exist as scripts, and neither README nor RESULTS reports a number from them. `permutation_auroc` and `holm` are defined and exported, no experiment calls them, and no p-value is reported. `results/` is gitignored and no script calls the record emitter, so no committed result file backs the tables; the numbers were copied by hand. The batched path has not been exercised against a real tokenizer. **Places where the repository disagrees with itself.** The README still says recall "cannot" be bounded; Theorem 8 supersedes that and RESULTS.md retracts it. The README header says "five proved bounds" while §4 says eleven theorems; Theorems 5 and 7 are both no-go results. RESULTS.md says Theorem 1 is loose by exactly one on the non-degenerate rows; the table gives 1, 1, 3, 2, 1, 4, 2 and 6. The `currency` Theorem 6 floor under `" the"`×128 is 0.917 in one docstring and in RESULTS.md and 0.833 in another. The README says there is no CI, and `.github/workflows/ci.yml` runs the test suite on push and pull request. It claims 163 tests; a static count at this commit finds 340 test functions, and the suite was not re-run for this essay. The README says `verdict.scores` includes every candidate, but candidates are skipped when the baseline floor is zero. It says a shared seed gives a shared projection matrix across models; the matrix's shape depends on the source width. The README's attribution lists four source repositories, and `governor.py` names a fifth, seal-demon-tts. **Open questions the repository names.** Whether the guard beats one-pass baselines at matched coverage; whether the Kaplan–Yorke dimension predicts block pruning; whether the transformer cocycle satisfies the hypotheses of Oseledets and Takens; whether `tail_alpha` is width-invariant; where the Krylov estimator overtakes the exact Jacobian; and whether `NEUTRAL_PREFIX` is neutral for any relation not tested. ## Read more [View the project](https://teerth.dev/caustic/) · [Source on GitHub](https://github.com/teerthsharma/caustic) The project site is [teerth.dev/caustic](https://teerth.dev/caustic/) and the source is [github.com/teerthsharma/caustic](https://github.com/teerthsharma/caustic). Every line reference in this essay is at commit [`6fed3df`](https://github.com/teerthsharma/caustic/tree/6fed3dfb630f854d64e5bdb1cd916d13d5abc8c2). Key files: - [`README.md`](https://github.com/teerthsharma/caustic/blob/6fed3dfb630f854d64e5bdb1cd916d13d5abc8c2/README.md): the measurements, the theorem statements and the limits, with the stale recall claim noted above. - [`RESULTS.md`](https://github.com/teerthsharma/caustic/blob/6fed3dfb630f854d64e5bdb1cd916d13d5abc8c2/RESULTS.md): every reported number with its control, the cross-model certificate table and the recall correction. - [`caustic/theorems.py`](https://github.com/teerthsharma/caustic/blob/6fed3dfb630f854d64e5bdb1cd916d13d5abc8c2/caustic/theorems.py): all eleven statements and proofs in the module docstring, with executable witnesses. - [`caustic/regime.py`](https://github.com/teerthsharma/caustic/blob/6fed3dfb630f854d64e5bdb1cd916d13d5abc8c2/caustic/regime.py), [`caustic/governor.py`](https://github.com/teerthsharma/caustic/blob/6fed3dfb630f854d64e5bdb1cd916d13d5abc8c2/caustic/governor.py), [`caustic/guard.py`](https://github.com/teerthsharma/caustic/blob/6fed3dfb630f854d64e5bdb1cd916d13d5abc8c2/caustic/guard.py): the partition, the governor and the inference-time guard. - [`tests/test_recall_no_go.py`](https://github.com/teerthsharma/caustic/blob/6fed3dfb630f854d64e5bdb1cd916d13d5abc8c2/tests/test_recall_no_go.py): the exhaustive Theorem 8 check. - [`notebooks/caustic_quickstart.ipynb`](https://github.com/teerthsharma/caustic/blob/6fed3dfb630f854d64e5bdb1cd916d13d5abc8c2/notebooks/caustic_quickstart.ipynb): the quickstart, including the 135M non-reproduction. Related essays on this blog: - [monodromy](/monodromy): the same question from the geometry side, whether a map can be run backwards, and why a pointwise Jacobian determinant check does not settle it. - [resolvent](/resolvent): another case where scoring each item on its own is the wrong rule, there for candidates in a world-model planner. - [sigmoid](/sigmoid): a world model built to say when it does not know, the abstention problem from another direction. --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # tangle > From a photo of two cables: an integer linking-number certificate, or a refusal that names the crossing to re-shoot. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/tangle - Repository: https://github.com/teerthsharma/tangle - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "tangle", 2026, https://teerth.blog/tangle - Languages: Python - Question: Can two cables in one picture be proved impossible to pull apart, with a refusal instead of a guess when the picture is not good enough? - Headline result: 0 wrong certificates over 2,000 closed-braid diagrams and 80 rendered scenes; 0 of 247 real photographs certified (control: abstain-on-any-unknown also 0 wrong; a coin flip on the same unknowns: 755 and 47 wrong) ## What it is Two cables are wound round each other behind a desk. Before pulling, there is one question worth answering: are they actually locked, or do they only look bad? Pull on a genuinely locked pair and you tighten it into a worse knot. I wanted a tool that answers that from a single photograph, and answers it in one of two ways only: a yes with a proof behind it, or a refusal with a reason. No maybe, and no percentage. tangle is that tool. It looks at a picture of two cables, builds the diagram of their crossings, and computes the linking number of the pair. The trick is in what it does with the crossings it cannot read. Every crossing has a sign, and the sign needs to know which cable is on top. When the photograph cannot say, tangle does not guess. It treats that crossing as unknown, and with $k$ unknown crossings and a signed sum $S$ over the readable ones, it computes the full set of linking numbers the photograph is consistent with: the interval $[(S-k)/2,\,(S+k)/2]$, in $O(k)$ steps, without enumerating the $2^k$ possibilities. From that interval there are three outcomes. If the interval excludes zero, the cables are certified LINKED: no motion separates them while their ends are held. If one cable is on top at every crossing between the two, a different witness certifies them SEPARABLE. Otherwise tangle refuses, and when the refusal comes from an unreadable crossing it names that crossing and the camera bearing to re-shoot it from. A linking number of zero is never read as "unlinked"; the Whitehead link has $\mathrm{lk} = 0$ and does not come apart. The reason for building it this way is a cost asymmetry, which the README puts better than I can paraphrase: "A refusal costs you a second photograph. A wrong certificate tells somebody to leave a live cable plugged in." The project splits cleanly in two. The certified layer takes a diagram and returns an exact verdict; its own docstring says the claim is "given a correct diagram, the verdict is exact", and that nothing in it is evidence that the diagram is correct. The imaging layer turns pixels into that diagram, and it is where the risk lives. I say this first because the measured results split the same way. On diagrams and on scenes tangle renders itself, it certifies often, and none of the certificates in the runs reported here was wrong. On real photographs it has certified nothing: 247 tries, 247 refusals, and not one diagram built. The figure below shows the first fact the design rests on. Swap which cable is on top at a crossing and the union outline of the two cables does not change by a single pixel. For opaque cables the silhouette carries no depth at all. The only place over/under shows up is in the colour of each cable: whether it runs continuously through the crossing or is interrupted there. **Figure 1.** Two scenes that are different tangles, the clasp and the same clasp with over and under swapped at every crossing, rendered with the repository's patch model. The left and middle panels show each scene; the right panel shows the union silhouette, which is identical for both, and the two difference maps below count the pixels that differ in the outline (none) and in colour (only near the crossings). Scene (clasp or stack), over/under swap, cable width and view are adjustable. The renderer has matte, constant-colour cables with no shadow, specular highlight, JPEG or lens, so this says nothing about photographs; it shows only that an outline cannot read depth. Colour key: structure: cable identity: A dark, B grey (no quantity); rings mark the crossings in the outline panel; ink: differing pixels in the two difference maps. ## What it can do The headline claim is narrow on purpose. It is not "tangle is accurate". It is that the $O(k)$ interval certifies strictly more pairs than refusing on any unknown crossing does, at the same zero error rate. The baseline is four lines of code: if any crossing between the two cables is unreadable, refuse; otherwise take the half-sum. That baseline also scores zero errors, so the gap between the two is the only number in this benchmark that could have come out badly. The corpus is 400 closed braids on two strands, six letters each, with $k = 0, 1, 2, 3, 4$ of the crossings between the two components set to UNKNOWN: 2,000 closed-braid diagrams. Ground truth is not the package's own arithmetic. For a closed two-strand braid, $\mathrm{lk}$ is half the signed exponent sum of the word, known in closed form before any code runs, and the bench asserts the diagram layer agrees with it on every entry. **Measured: 23.6%** certified at 0 wrong, on 2,000 closed-braid diagrams with injected unknown crossings (≤0.6% of certified at 95%). Control: abstain on any unknown crossing: 13.2% certified, 0 wrong (≤1.1% of certified at 95%); gap printed as 10.3 points. n = 400 braid seeds × k = 0..4. Source: bench.py, RESULTS.md:45-55 @ 563732b; measured at ac84ca0, WIN-16QAL06O9GB, Python 3.11.9. (Note: The two printed percentages subtract to 10.4. `bench.py` computes the gap from the entry counts before rounding, which gives 10.3.) Zero wrong is the expected result here, and I want to be plain about why. The corpus corrupts the input in exactly the one way the interval theorem is proved against, so a correct implementation cannot produce a wrong certificate on it. The `0 wrong` row measures the code, not the design, and it ships as a regression test. The control that shows the corpus is not trivially easy is to keep the same diagrams and replace the decision rule with a coin flip on every unknown crossing. **Measured: 755** wrong certificates when the unknown crossings are resolved by a coin flip instead of the interval. Control: tangle on the identical diagrams: 0 wrong. n = 2,000 entries; coin flip certifies 70.3%. Source: RESULTS.md:54-55 @ 563732b; measured at ac84ca0. Of the 2,000 entries tangle certified 23.6%, returned NOT CERTIFIED on 6.2% and refused 70.2%. The refusal rate is a function of the blur schedule, up to four of six crossings erased on purpose, and is not a photograph's refuse rate. Among certified verdicts the number of unknown crossings was $k = 0$ for 276, $k = 1$ for 147 and $k = 2$ for 49; nothing above $k = 2$, because certification needs $|S| > k$ and six letters run out. The second benchmark goes from picture to verdict. 20 seeded piles of two cables, each rendered under four nuisance arms (clean, blur 1.0 px, blur 3.0 px, antialiased), give 80 scenes. Truth comes from the scene's own height functions, so over/under is never authored: this scores the reader against the scene rather than against itself. **Measured: 63 certified, 0 wrong** 80 rendered scenes, picture to verdict; rule of three gives an upper bound of 0.048, not a rate of 0. Control: coin flip in place of the over/under reader on the identical extracted diagrams: 47 wrong certificates. n = 20 piles × 4 arms; per arm certified/wrong/refused 16/0/4, 17/0/3, 16/0/4, 14/0/6. Source: tests/test_vision.py, RESULTS.md:142-160 @ 563732b; measured at ac84ca0. On the crossings the tracer accepted, over/under was read the right way round 182 times out of 182 (46, 47, 44 and 45 per arm). That says nothing about the crossings it refused. Then the real images. I collected 247 free-licensed pictures the repository did not make: 99 rope and knot photographs and 72 cabling photographs from Wikimedia Commons, 27 published link diagrams, and 49 single-curve drawings from a Hugging Face set, 51 MB in all, each pinned by sha256. **Measured: 0 of 247** real images certified; 0 wrong certificates of 19 labelled; 0 diagrams built at all. Control: the same harness on 20 rendered piles written to PNG and read back: 13 certified (65.0%). n = 247 images: 104 BRANCHED_SKELETON, 65 NO_INTENSITY_GAP, 61 NOT_TWO_COMPONENTS, 17 OPEN_TRACE. Source: real.py, RESULTS.md:233-276 @ 563732b; fetched 2026-09-05, measured at ac84ca0. Not one real image reached the certifier. Every one was stopped by a precondition of the tracer before a diagram existed, and the most common stop, at 104 of 247, was `BRANCHED_SKELETON`: a cable's centreline forks because it crosses itself or because two strands merge into one blob. The repository attributes the zero to the corpus rather than the plumbing, because 13 of 20 rendered piles certify through the identical code path, same resize, same alpha handling, same entry point. I think that reading is right, and I also think it should be said without softening: today tangle answers on pictures it renders itself, and on real photographs it says it cannot see, every time. The 19 link diagrams that carry a published linking number score right 0, WRONG 0, no verdict 19. The WRONG column is empty because the tool never answered. What the certified layer does on a diagram, end to end through the CLI, looks like this. The clasp scene certifies; the pile refuses and says where to stand: ```text python -m tangle clasp.png # synth.clasp(sign=+1), seed 1 CERTIFIED LINKED lk = 1 interval [1, 1] S = 2 k = 0 advice lk = 1 over all 2^0 resolutions. Unplug one. exit 0 python -m tangle pile5.png # synth.pile(5), seed 5 REFUSED LK_STRADDLES_ZERO look at crossing 1 at 275,279, camera bearing 152 deg advice the achievable interval [0, 1] contains 0 exit 2 ``` The truth for that pile is $\mathrm{lk} = 0$, and the tool does not say "unlinked"; it says the achievable interval is $[0, 1]$. Exit codes are the verdict, so a shell can branch on them: 0 certified, 1 not certified, 2 refused, 3 bad input. The next figure puts the same certified layer in your hands: draw two cables, set over/under at each crossing or leave it unknown, and read the verdict tangle would print. **Figure 2.** The certified layer end to end on a diagram you draw. Click to place the vertices of cable A and cable B; crossings are found by segment intersection, each crossing cycles through A over, B over and unknown, and the verdict card is computed by a port of certify(): interval, S, k, the crossing to look at and its camera bearing. Over/under here is typed by you, not read from pixels; the image tracer's results on real images are in the refusal-wall figure under Limitations. Only the absolute value of lk is certified, and image coordinates are y-down. Colour key: structure: cable identity: A dark, B grey; proof: CERTIFIED: interval excludes 0, or the over-everywhere witness fires; refused: REFUSED or unknown crossing (hatched). ## How it was made ### The object The object is not a closed link. A photograph shows two pieces of cable that leave the frame, and closing them up at the frame boundary was an earlier design; it was wrong, because the exit points move with the camera, and I deleted the closure. What remains is a two-string tangle with pinned ends. > **Definition: Tangle linking number with pinned ends.** > > The scene is a ball. Each cable is an arc properly embedded in it, with its two ends fixed where the cable leaves the picture. The verdict is invariant under any motion of the cables inside the scene that keeps them disjoint and keeps all four ends fixed. Nothing is claimed about what the cables do outside the picture. The invariant is the linking number of the two arcs, computed as half the sum of the signs of the crossings between them. That convention is a string constant in `certify.py` and travels on every verdict. ### One crossing, two factors The modelling decision the whole certificate rests on is that a crossing's sign factors into two parts. One part is planar: the sign of the determinant of the two cables' in-plane tangent directions at the crossing. It needs no depth at all. The other part is the over/under bit, the only thing a photograph might fail to read. $$ \varepsilon(c)\;=\;\underbrace{\operatorname{sgn}\det\big[\,t_a\;\;t_b\,\big]}_{\mathrm{base}(c)\,:\ \text{in-plane, no depth}}\;\cdot\;\underbrace{x_c}_{\pm 1,\ +1\ \text{iff}\ A\ \text{is over}} $$ **Figure 3.** One crossing. The tangent of cable A is fixed and the tangent of cable B is rotated by the reader; the determinant of the two tangents is drawn as a vector perpendicular to the plane, pointing up or down with its sign, so base(c) is literally which way the arrow points. The over/under choice sets x_c, and the product is the crossing sign. When |sin θ| < 0.15 (the tangents nearly parallel), base is UNKNOWN: that widens the interval and never refuses. The eight small multiples sweep the angle in 22.5 degree steps. Heights are schematic. Colour key: structure: tangents t_a (dark) and t_b (grey); parallelogram sign by pattern and arrow direction; refused: tangents too nearly parallel: base UNKNOWN (hatched wedge). In the code, `base` is the sign of the 2D cross product of the two tangents, set to unknown below `SIN_MIN = 0.15`, and the sign is the product, unknown if either factor is: ```python # tangle/diagram.py:136-140 @ 563732b def sign(self, c: Crossing) -> int | None: """base * (+1 if a is over else -1); None if either factor is unknown.""" if c.base is None or c.over is None: return None return c.base if c.over == "a" else -c.base ``` Image coordinates are y-down throughout, which negates every crossing sign relative to the usual mathematical convention. Only $|\mathrm{lk}|$ is certified and the sign is a stated convention, so this costs nothing, but it is written into the module docstring so nobody re-derives it from a failing test. ### The half-sum The linking number is half the sum of the signs over the set $C(A,B)$ of crossings between the two cables. Self-crossings do not enter it. $$ \mathrm{lk}(A,B)\;=\;\frac{1}{2}\sum_{c\,\in\,C(A,B)}\varepsilon(c) $$ **Figure 4.** A two-cable closed braid built by the port of Diagram.from_braid. Click a crossing to flip its letter; every sign is recomputed from the real cross product of the tangents, and the readout checks the half-sum against the closed form, half the signed exponent sum of the word. Inserting a cancelling pair of crossings (a Reidemeister II move) changes the crossing count and leaves lk unchanged; mirroring negates it. This is diagram-move invariance, which is weaker than camera invariance, and the T(2,4) preset with |lk| = 2 is a word the rendered-scene corpus never produces. Colour key: structure: component identity: dark and grey; filled disc = +1, hollow ring = −1; proof: lk ≠ 0: a LINKED certificate exists; refused: no LINKED certificate: NOT CERTIFIED or SEPARABLE verdict. The half-sum is an isotopy invariant: bend the cables, drape them, re-route them, shoot from the other side, and as long as the four ends stay pinned where they leave the picture, $\mathrm{lk}$ does not move. That is the property a network trained on rope photographs does not have, and it is why I built on an invariant rather than a classifier. None of this layer is new; the half-sum goes back to Gauss in 1833. ### From pixels to a diagram The imaging layer is a chain of stages, and every stage is terminal on refusal. In order: estimate the background from a 12 px border ring in Lab colour; threshold the cable-ness image with Otsu and refuse if the two classes are not separated; clean up the mask; split it into two cables with k-means on colour, refusing if one cluster is empty, too small, or within a just-noticeable difference of the other; estimate the cable width from the distance transform; thin each cable to a skeleton and prune spurs; refuse with `BRANCHED_SKELETON` if any skeleton pixel still has four or more neighbours; chain the pieces of each cable across occlusion gaps into one curve from frame edge to frame edge, or back to its start, refusing with `OPEN_TRACE` otherwise; pin the ends to the frame; and build the crossings with the same `Diagram.from_polylines` the closed-form tests use, so crossing numbering, `base` and angle are shared code. The segmentation gate deserves its equation, because it decides most of what a photograph can do. Otsu always returns a threshold, so the refusal comes from Fisher's ratio of the two classes it found: $$ F\;=\;\frac{\mu_{\mathrm{hi}}-\mu_{\mathrm{lo}}}{s_{\mathrm{hi}}+s_{\mathrm{lo}}},\qquad \text{refuse if } F < 2.0 $$ **Measured: 182 of 247** real images admitted by the Otsu-plus-Fisher gate (106 of the 171 photographs). Control: the earlier widest-empty-histogram-run rule on the same images: 97 of 247 (21 of 171). n = 247 images, one pass; F = 2.0 calibrated against no-object distributions: Gaussian 1.33, half-normal 1.46, uniform 1.73. Source: RESULTS.md:315-324 @ 563732b; measured at ac84ca0. The earlier rule required a wide empty run in the histogram of cable-ness. That is satisfiable only by a renderer: a rendered cable is a flat stroke on a flat field, so its histogram is two spikes with nothing between, while a photograph's is dense in every bin. On the rendered side the new gate also widened the noise envelope from $\sigma = 6/255$ to $\sigma = 16/255$, with 0 unsound certificates across a seven-level, twenty-pile sweep (a figure the repository states in a test docstring). ### Reading over and under Over/under is read from occlusion. Where one cable passes under the other, its own mask is interrupted, and the tracer has to bridge that gap to chain the cable into one curve. The bridged strand is the one underneath. A bridge that explains a crossing at angle $\theta$ should be about $w/\sin\theta$ long, for cable width $w$, so the confidence of a reading is the product of a margin term (one strand bridged, the other not) and a fit term (the bridge has the length its own crossing angle predicts): $$ \mathrm{conf}\;=\;\underbrace{\frac{|g_a-g_b|}{g_a+g_b}}_{\text{margin}}\cdot\underbrace{\exp\!\left[-\left(\frac{\log\!\big(\max(g_a,g_b)\,/\,\mathrm{pred}\big)}{S}\right)^{2}\right]}_{\text{fit}}, \qquad \mathrm{pred}=\frac{K\,w}{\sin\theta} $$ **Figure 5.** The over/under reader on one crossing. The under strand is drawn with a bridge of adjustable length, the over strand with an optional second break (contradicting evidence); the dimension lines on the under strand mark the over strand's footprint w/sin θ and the bridge length g, and the bridge the angle predicts is K times the footprint, K·w/sin θ. The gauge shows margin, fit and their product, against the gate TAU = 0.80, below which the certified layer downgrades the reading to UNKNOWN. Panel C sweeps the same formula over crossing angle and bridge ratio, coloured by conf on the sequential scale under the map. conf is a score, not a probability; K = 1.165 was fitted on the repository's own matte renders and has never been measured on a photograph, and this draws the model's bridge, not a measured one. Colour key: structure: cable identity: a dark, b grey; proof: conf ≥ TAU: reading kept, the certified layer uses it; refused: conf below TAU: downgraded to UNKNOWN (hatched); seq: Panel C: conf, 0 (light) to 1 (dark), scale under the map. Here $g_a, g_b$ are the summed bridge lengths near the crossing on each cable, $K = 1.165$ and $S = 0.85$. The repository reports that $K$ was measured over 72 bridges on 36 seeded piles, with a spread of 0.132 in log units; $S$ is set wide on purpose, to catch a bridge twice or half its predicted length. The code is a direct transcription: ```python # tangle/vision.py:610-620 @ 563732b for c in d.crossings: ga = _gap_near(bridges[c.a.cable], c.xy, GAP_R_W * w_est) gb = _gap_near(bridges[c.b.cable], c.xy, GAP_R_W * w_est) if max(ga, gb) < NOISE_W * w_est: out.append(replace(c, over=None, over_conf=0.0, kind="unknown")) continue over = "b" if ga > gb else "a" margin = abs(ga - gb) / (ga + gb) pred = BRIDGE_K * w_est / math.sin(math.radians(c.angle_deg)) fit = math.exp(-((math.log(max(ga, gb) / pred) / BRIDGE_S) ** 2)) out.append(replace(c, over=over, over_conf=float(margin * fit), kind="read")) ``` A crossing where neither strand broke has no evidence and is UNKNOWN with confidence 0, never a low-confidence reading. The gate `TAU = 0.80` lives inside the certified layer, not in the caller: `certify()` downgrades any reading below it to UNKNOWN itself, so the honesty boundary sits where a test can hold it. It is a per-rig calibration knob, not a constant. ### The certified layer, in order `certify()` runs a fixed sequence, and every refusal is terminal. It checks for defects (a free cable end inside the frame, two crossings at one point); then whether the four exit points interleave on the frame, which makes $\mathrm{lk}$ a half-integer; then a parity guard; then the interval; then, only if the interval did not certify, the over-everywhere witness; then a refusal that names a crossing if any crossing is unknown; and otherwise NOT CERTIFIED, either because both cables are on top somewhere (it names the minority crossings, which are the obstructions) or because $\mathrm{lk} = 0$, which is not evidence of anything. The parity guard is free. For two arcs whose ends do not interleave, $S + k$ must be even. A tracer that misses or invents a single crossing between the two cables makes it odd, and the verdict is `ODD_CROSSING_PARITY`. For the case the theorem covers, two arcs with ends on the frame or two closed components, that catches every odd-sized tracer error. A pair of one arc and one closed curve is outside it, and there the guard can refuse a correct diagram; that failure is a refusal, never a certificate. The guard cannot catch an even number, which is the subject of the Limitations section. The repository reports 226 tests passed and 2 skipped. I did not re-run them for this essay; the count is the repository's own. ## What's new in it The usual way to answer "is this rope knotted" from a photograph is a learned classifier: show a network many pictures and have it predict a class or a crossing number. The usual way to make a classifier safe is to let it abstain when its confidence is low. And the usual way to handle a crossing whose over/under cannot be read is either to drop the whole input or to commit to a best guess. tangle replaces all three with one exact statement. Both factors of $\varepsilon(c)$ are $\pm 1$, so the linking number is affine in every crossing the photograph could not read. Each unknown crossing contributes either $+1$ or $-1$ to the sum, independently. With $S$ the signed sum over the readable crossings between the two cables and $k$ unreadable ones, the set of linking numbers achievable over all $2^k$ resolutions is exactly $k+1$ consecutive integers: $$ \Big\{\,\mathrm{lk}\,\Big\}_{2^k}\;=\;\Big\{\ \tfrac{S-k}{2}+j\ :\ j=0,1,\dots,k\ \Big\}, \qquad \mathrm{lk}_{\min}=\tfrac{S-k}{2},\qquad \mathrm{lk}_{\max}=\tfrac{S+k}{2} $$ **Figure 6.** The interval theorem, run live. Six crossings between two cables, each set to +1, −1 or unknown. The bar on the number line is \[(S−k)/2, (S+k)/2]; it is a certificate when it clears the zero bar and a refusal when it touches it. Below, all 2^k resolutions of the unknown crossings are enumerated independently and their linking numbers plotted on the same axis, so the reader can check that the enumeration lands on exactly the integers the formula predicts; the column heights count lifts and are not probabilities. Resolving any unknown crossing narrows the bar by exactly one. The tracer-error button deletes one crossing and the parity guard refuses. The comparator strip shows what abstain-on-any-unknown returns on the same input. Colour key: proof: interval excludes 0: CERTIFIED LINKED; refused: interval straddles 0: REFUSED (hatched); hatched square = unknown crossing; baseline: abstain-on-any-unknown verdict on the same crossings; structure: crossing sign glyphs: filled disc = +1, hollow ring = −1. **Measured: 0 / 1000** random patterns on which the O(k) interval disagrees with explicit 2^k enumeration. Control: brute_force_interval, which resolves every lift through Diagram.resolve and Diagram.sign and shares no arithmetic with the formula. n = 192,540 lifts enumerated, k ≤ 10, seed 20260905. Source: tests/test_certify.py, RESULTS.md:100-110 @ 563732b; measured at ac84ca0. The code is short enough to quote whole: ```python # tangle/certify.py:122-139 @ 563732b def lk_interval(d: Diagram, i: int = 0, j: int = 1) -> Interval: """T4, in O(k), with no enumeration. S is the signed sum over the readable inter-component crossings; k counts the ones whose sign is unknown for any reason -- unreadable over/under, or a tangent too nearly parallel for base. Both are the same unknown +/-1 in the same product. """ S = 0 k = 0 for c in d.between(i, j): s = d.sign(c) if s is None: k += 1 else: S += s if (S + k) % 2 != 0: raise ValueError(f"S + k is odd (S={S}, k={k}); parity_ok must be checked first (T2)") return Interval(lo=(S - k) // 2, hi=(S + k) // 2, known_sum=S, unknown=k) ``` Three things follow, and they are what I consider new in this project as opposed to in its mathematics. First, the decision rule has no fourth case. Certification is $\mathrm{lk}_{\min} > 0$ or $\mathrm{lk}_{\max} < 0$, and because the interval is contiguous with step 1, those two plus "straddles zero" are exhaustive. This is an interval certificate, not a confidence score. Flipping the over/under at any crossing of a plane diagram yields another diagram a real pair of cables could form, so the $2^k$ orbit is the exact set of tangles consistent with what the camera saw, not a superset and not a sample. Second, the certificate points one way. A nonzero linking number proves the cables cannot be separated with their ends held; a zero linking number proves nothing. Separability comes from a different witness that points the other way: if one cable is the over-strand at every crossing between the two, and the ends do not interleave, a disk separates them. That witness needs every one of those crossings read, and it never fires from $\mathrm{lk} = 0$. The word *unlinked* is a banned substring in the package, enforced by a test that greps the source and every verdict the example set can produce. Third, a refusal is an answer with an instruction. When the interval straddles zero, tangle computes how many crossings must be resolved before any certificate is possible, $r_{\min} = (k - |S|)/2 + 1$ (necessary, not sufficient), and names the unknown crossing with the most tangential angle, together with a camera bearing along the bisector of its two tangents. I should be exact about how much that naming is worth. The ranking is a perception heuristic about which crossing a new view is most likely to resolve. It is not an information criterion: because $\mathrm{lk}$ is affine, every unknown shrinks the interval by exactly one and no crossing is more decisive than another. I predicted before measuring that it would not beat random, and it did not: 19.9% certified after re-shooting the named crossing, against 20.0% for a uniformly random crossing at the same budget, over 1,404 straddles. ## What no one else built The certified layer is classical mathematics and I make no claim on it. The question for this section is whether the combination around it, an exact interval over the unreadable crossings of a photographed diagram used as a decision rule with refusals, exists elsewhere. These are the closest pieces of prior work I could find, and how each differs. **Matsuno, Tamaki, Arai and Fukuda (2006)**, [*Manipulation of deformable linear objects using knot invariants to classify the object condition based on image sensor information*](https://doi.org/10.1109/tmech.2006.878557), IEEE/ASME Transactions on Mechatronics. By its title this reads as the same pipeline: image, topological rope model, knot invariant, decision, twenty years ago. The paper is paywalled and I have not read it. If its invariant turns out to be the linking number, my claim scopes down to the explicit UNKNOWN state per crossing, the interval over the unknown-crossing orbit, the pinned-end tangle, and the refusal that names a crossing. **Dranowski, Kabkov and Tubbenhauer (2025)**, [*On knot detection via picture recognition*](https://arxiv.org/abs/2510.06284). Their stated goal is to take a photo of a knot and have a phone recognise it; the present work gives CNN and transformer baselines that predict crossing number directly from images, with a planned route through planar-diagram codes to invariants. Their abstract does not describe a way to carry unreadable crossings through to the invariant, or to abstain. tangle predicts nothing: it computes the set of values the readable crossings allow and answers only when that set decides. **Knots-10**, Nie and Yue, [*Physical Knot Classification Beyond Accuracy*](https://arxiv.org/abs/2603.23286). A 1,440-image, ten-class benchmark of physical knots, trained on loose knots and tested on tight ones, reporting that phone-photo accuracy drops by 58 to 69 percentage points. It has what tangle lacks, real rope photographs at scale. tangle has no training at all, and the invariant does not depend on material or colour by construction. I have not run tangle on their images, so I do not claim it does better there; on the 247 real images I did run, it answered nothing. **HANDLOOM**, Viswanath et al., [*Learned Tracing of One-Dimensional Objects for Inspection and Manipulation*](https://arxiv.org/abs/2303.08975). A learned tracer that fits a trace to a greyscale image of cables and classifies crossings, trained on simulated and real examples and used for inspection and robot manipulation. It is a better tracer than mine on real images. What tangle adds is downstream of any tracer: an exact integer from the trace, and a refusal when the trace does not decide it. **KnotDLO**, Dinkel et al., [*Toward Interpretable Knot Tying*](https://arxiv.org/abs/2506.22176). It acts: it ties an overhand knot with one hand, with no demonstrations or training, succeeding in half of 16 trials from unseen configurations. tangle does not act. It certifies a topological fact about a still picture. **SnapPy and Spherogram** ([docs](https://snappy.computop.org/spherogram.html)) and **pyknotid** ([docs](https://pyknotid.readthedocs.io/)). These compute every invariant tangle computes, and many more, better and more generally. Spherogram builds a link from a planar-diagram code in which every crossing is specified; pyknotid works from space curves, coordinates that already contain depth. Both start from an input in which over/under is known. tangle's whole contribution sits in the step they assume away: getting that input from a photograph, and saying exactly what follows when part of it is missing. They are also the third-party cross-check tangle does not yet have. **KnotPlot**, Scharein ([knotplot.com](https://knotplot.com/)). An interactive program for visualising and building three- and four-dimensional knots, from a database, by sketching, or by construction. Its input is a curve the user makes, not a photograph. **Selective classification.** [Chow (1970)](https://doi.org/10.1109/TIT.1970.1054406) set out the trade-off between error and rejection, and [Geifman and El-Yaniv (2017)](https://arxiv.org/abs/1705.08500) build a selective classifier on a trained network by thresholding its softmax response, so that a chosen error rate holds with high probability at test time. tangle shares the instinct, refuse rather than err, but the mechanism is different. Its refusal is not a threshold on a score; it is the outcome when the exact set of consistent answers contains zero. Given a correct diagram there is no error rate to bound. I should not overstate this: the over/under reader upstream does use a thresholded score, `TAU`, and that is exactly where the guarantee stops. Put together: in the work I could find and read, I did not find a system that keeps an explicit UNKNOWN per crossing of a photographed diagram, computes the exact set of linking numbers over all resolutions of those unknowns in $O(k)$, certifies only when that set excludes zero, certifies separability through an independent witness, and otherwise refuses and names the crossing to re-shoot. With Matsuno et al. unread, I hold that as a statement about what I found, not a claim to be first. The two figures below run the comparison that matters: the same input through three decision rules, live and then as measured. **Figure 7.** The same 400 closed two-strand braids, with k crossings blurred, run through three decision rules on identical input: tangle's interval, abstain on any unknown crossing, and a coin flip on each unknown crossing followed by the same certifier. The figure enumerates every blur choice and every coin outcome exactly, and stacks each rule's verdicts into certified-and-right, certified-and-wrong and not certified. Below, a single diagram shows each rule's verdict next to the truth. The coin stands for an over/under reader with no notion of uncertainty, not for any published method. The repository's single seeded draw can be overlaid. Colour key: proof: certified and right; baseline: certified and wrong (the old path: guessing the unknowns); refused: not certified or refused (hatched); ink: the repository’s single seeded draw (optional overlay). **Figure 8.** The measured headline. Top: on the 2,000 closed-braid diagrams, tangle's verdicts by number of unknown crossings k = 0..4, beside the abstain-on-any-unknown baseline, which certifies only at k = 0; the repository's draw of 276, 147 and 49 certified is marked. First table under the stage: tangle 23.6% at 0 wrong, abstain 13.2% at 0 wrong, coin flip 70.3% with 755 wrong. Second table: the 80 rendered scenes per nuisance arm, with the coin flip's 47 wrong certificates. On rendered scenes every certified verdict had k = 0, so there the interval certified nothing the plain half-sum would not have; the 70.2% refused on braids is set by the blur schedule and is not a photograph's refuse rate. Colour key: proof: tangle CERTIFIED; baseline: abstain baseline and coin flip; coin-flip wrong certificates; refused: REFUSED (hatched); structure: NOT CERTIFIED (outline only); ink: the repository’s single seeded draw (optional ticks). ## Limitations The first limitation is the one already in the hero. **tangle has been run on real images and it certified none of them.** 247 free-licensed images produced 0 certified verdicts, 0 wrong certificates, and 0 diagrams built. The dominant refusal is `BRANCHED_SKELETON`, 104 of 247: on a real pile most crossings are self-crossings, a knot photograph is one rope crossing itself, and a published link diagram is drawn with a black outline that welds strands into one region. The repository attributes the zero to the corpus, not the plumbing, because the same harness certifies 13 of 20 rendered piles through the same PNG round trip. Both halves of that are true at once: the harness works, and the corpus of real pictures is outside what it can trace. Resolving a forked centreline by continuing each cable's direction through the junction, with a margin so an ambiguous blob still refuses, is the one change that could move the real number off zero. It is not built. **Figure 9.** Panel A: all 247 real images by refusal reason and source (99 rope photographs, 72 cabling photographs, 27 link diagrams, 49 single-curve drawings that serve as an out-of-domain control and are all refused NOT_TWO_COMPONENTS), with 0 certified and 0 wrong. Every segment is refused; its count is printed where the segment is wide enough. Below it, the control strip of 20 rendered piles through the identical harness, 13 certified and 7 refused, drawn at the same width per image, so the scale bar reads for both. Panel B: the Otsu-plus-Fisher gate computed live on a model histogram of cable-ness, with separation and cable share adjustable. The noise slider rescales the histogram; F depends on separation and share, not on the noise level. The refusal line is F = 2.0. The two measured noise-cliff markers (16/255 traced 10 of 10, 26/255 traced 0 of 10) are the repository's results and are not computed by the model. Panel B models the gate, not a photograph. Colour key: withdrawn: refused (hatched), segmented by reason; proof: certified (control strip), or admitted by the gate at F ≥ 2; ink-3: cable-ness histogram, counts (no hue); ink: Otsu threshold; measured noise-cliff markers; structure: class means (ticks); parameter: the noise you set. The rest, stated flatly: - **Two visually distinct cables, or it refuses.** Colour does segmentation, so a pile of identical black charging cables is one mask and `NOT_TWO_COMPONENTS`: 61 of 247 real images. The most common real scene is the worst case. - **Noise is a cliff, not a slope.** At $\sigma = 16/255$ ten of ten rendered piles trace; at $26/255$ none do, all refused `NO_INTENSITY_GAP`. The failure direction is always a refusal. - **Every benchmark percentage except the real-image table comes from the repository's own renderer or from closed-form diagrams.** That includes the coverage table, the 182 of 182 over/under readings, the 47 wrong coin-flip certificates, `TAU`, `BRIDGE_K`, and the blur and antialiasing arms. The renderer draws matte, constant-colour cables with no shadow, specular highlight, JPEG or lens, and rejects self-crossings, crossings closer than four widths and crossing angles below 25 degrees. A renderer cannot falsify the module that reads it. - **On rendered scenes the interval theorem certified nothing the plain half-sum would not have.** All 63 certified verdicts had $k = 0$. The 10.3-point gap exists on the braid corpus, where the unknowns are injected on purpose. - **No camera model.** Invariance is tested as diagram-move invariance (Reidemeister moves), which is weaker than camera invariance. The two-view interval intersection exists in `certify.intersect`, where disjoint intervals would prove one trace wrong, and it is not wired up. - **The rendered corpus cannot exhibit $|\mathrm{lk}| \ge 2$.** An arch weaving across an arch cannot wrap twice; $|\mathrm{lk}| \ge 2$ lives only in the torus-link family, which never goes through a camera. - **Only $|\mathrm{lk}|$ is certified.** The sign is a stated convention, because image coordinates are y-down. - **The certificate answers a narrower question than you will ask.** "Cannot be separated with the ends held" is not "will be annoying to untangle". A pile with $\mathrm{lk} = 0$ can still be a nightmare of friction. - **The Alexander determinant is not part of the verdict.** `alexander.py` computes $\det(L) = |\Delta_L(-1)|$ from a Goeritz matrix, is exported, and is tested against a closed-form ladder ($\det T(2,n) = n$, figure-eight 5, Whitehead 8). But `certify.py` and the CLI do not call it; the verdict is the linking number only. The determinant is nonlinear in the unknowns, so no interval of the $O(k)$ kind exists for it, and it would need the $2^k$ enumeration under a budget of $2^{16}$ lifts. - **Two controls named in the specification were not run**: the R2-drape alternation control and the mask-overlap control. `bench.py` prints their absence. The one limitation I consider most important is quieter than the zero. **An even number of tracer errors is not caught**, and it can produce a confidently wrong certified integer. Parity catches every odd-sized error for free; nothing in a single view catches a pair. The repository calls this the live false-certification path. The figure below shows it directly: delete crossings from correct diagrams and watch the guard refuse at odd counts and pass at even ones, sometimes with the wrong integer. **Figure 10.** The parity guard against injected tracer errors on the same 400 closed braids. Top row: the tracer misses m of the six crossings between the two cables, for m = 0..4. At odd m the guard refuses every diagram; at even m it passes them all, and some of those diagrams certify a linking number the truth contradicts. Bottom row: over/under misread at j crossings while staying 'known', which never changes the parity and is never caught. Counts are computed exactly by the figure on the repository's braid corpus under this stated error model; they are not rates on photographs, where no image reached the certifier. Colour key: withdrawn: refused by parity (hatched); ink: wrong certificate (solid ink with hatch): the limitation; proof: certified and right. The design itself went through several verdicts that are now dead. They are listed with what killed each, because each one changed the code. **What failed.** - Withdrawn: "lk = 0 means just pull." Killed by: The Whitehead link has lk = 0 and does not come apart. The word unlinked is now banned in code (certify.py:79).. - Withdrawn: "Closing the diagram at the frame boundary makes lk camera-invariant." Killed by: The exit points move with the camera. The closure was deleted; the object is a 2-string tangle with pinned ends.. - Withdrawn: "A width bump at the crossing reads over/under." Killed by: For opaque cables the silhouette is provably depth-blind; the cue was reading the renderer's drop shadow. Test: test_silhouette_carries_no_depth.. - Withdrawn: "A contraction radius fixed in cable widths merges the skeleton's H-pattern." Killed by: The bridge at crossing angle θ is about w / sin θ long, so a constant radius loses shallow crossings. Replaced by the angle-aware rule, vision.BRIDGE_K.. - Withdrawn: "det = 5 certifies a figure-eight tie-in." Killed by: A follow-through is tied on a bight, so the traced curve is the unknot, det = 1. No code path turns a determinant into a knot name.. - Withdrawn: "Threshold at the widest empty gap in the histogram." Killed by: Satisfiable only by a renderer. On the same 247 real images it admitted 97; Otsu plus Fisher admits 182.. - Withdrawn: Naming the crossing to re-shoot beats re-shooting a random one. Killed by: 19.9% against 20.0% over 1,404 straddles; lk is affine, so every unknown shrinks the interval by exactly 1.. - Withdrawn: The interval theorem adds coverage on rendered scenes. Killed by: All 63 certified rendered verdicts had k = 0.. Never claimed, in any code path or string: unlinked from $\mathrm{lk} = 0$; unknotted from $\det = 1$; a knot name from any determinant; chirality; a probability or percentage attached to a verdict; any climbing, rigging or safety verdict; anything about the cables outside the frame. ## Read more [View the project](https://github.com/teerthsharma/tangle) · [Source on GitHub](https://github.com/teerthsharma/tangle) tangle has no project site of its own; the repository is the place to go. It installs from PyPI as [`tanglekit`](https://pypi.org/project/tanglekit/); the import and the command stay `tangle`. Short link: [teerth.dev/tangle](https://teerth.dev/tangle). Every file reference in this essay is pinned at commit `563732b`. - [README.md](https://github.com/teerthsharma/tangle/blob/563732b288e2ffd2d9aed1f0a90b31cc3b0d3801/README.md): the statement, the verdict table, the benchmarks, prior art, what was wrong, and the limits collected once. - [RESULTS.md](https://github.com/teerthsharma/tangle/blob/563732b288e2ffd2d9aed1f0a90b31cc3b0d3801/RESULTS.md): every table with the command that produced it, every control, every arm that lost, and the claims not earned. - [tangle/certify.py](https://github.com/teerthsharma/tangle/blob/563732b288e2ffd2d9aed1f0a90b31cc3b0d3801/tangle/certify.py): the certified layer, theorems T2 to T7, the interval, the two witnesses and the refusals. - [tangle/diagram.py](https://github.com/teerthsharma/tangle/blob/563732b288e2ffd2d9aed1f0a90b31cc3b0d3801/tangle/diagram.py) and [tangle/vision.py](https://github.com/teerthsharma/tangle/blob/563732b288e2ffd2d9aed1f0a90b31cc3b0d3801/tangle/vision.py): the sign factorisation, and the imaging layer with its refusals. - [bench.py](https://github.com/teerthsharma/tangle/blob/563732b288e2ffd2d9aed1f0a90b31cc3b0d3801/bench.py) and [real.py](https://github.com/teerthsharma/tangle/blob/563732b288e2ffd2d9aed1f0a90b31cc3b0d3801/real.py): the coverage benchmark and the real-image run. Related essays on this blog: [nerve](/nerve), which computes linking numbers on polymers; and the other two certify-or-refuse tools, [separatrix](/separatrix) and [planimeter](/planimeter). --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # separatrix > Tells you whether a top-k list came from your data or from the computer's rounding, and refuses, naming the pair, when it cannot tell. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/separatrix - Repository: https://github.com/teerthsharma/separatrix - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "separatrix", 2026, https://teerth.blog/separatrix - Languages: Python - Question: Did my data choose this top-k set, or did floating-point rounding choose it? - Headline result: 0 of 1,116 certified top-10 sets moved across nine numerically distinct engines (control: all 8 sets that did move had been refused first (one draw, seed 11)) ## What it is You ask a nearest-neighbour search for the ten closest items and you get ten back. Most of them are not in doubt. The last one or two, though, can sit so close to the eleventh that which one made the list was, in the words I used in the README, "settled by the computer's rounding rather than by your data". That is why the same query over the same data can hand back a different tenth result on another machine, or after somebody changes a batch size. separatrix is a small Python package that answers one question about such a list. It returns a top-k set, an argmin, or a threshold decision only when a rigorous a-priori bound on the rounding error proves that rounding could not have changed it; the verdict is then `CERTIFIED`. When it cannot prove that, it returns a typed `REFUSED` verdict that names the boundary pair, prints both of their intervals, the gap between them and how wide the intervals are, and says what to change. It does not guess, and it does not quietly pick one of the two. separatrix is a tool that came out of [Epsilon-Hollow](/epsilon-hollow). The people who hit the underlying problem are anyone whose distance code goes through the Gram identity. `torch.cdist` switches to it above 25 rows because it is faster, and hand-rolled Gram code uses it as well. It turns a distance computation into one matrix multiply: $$ d^2(x, y) \;=\; \lVert x \rVert^2 + \lVert y \rVert^2 - 2\,\langle x, y \rangle . $$ **Figure 1.** Two distinct float64 points, x = (M, 0) and y = (M + delta, 0). The figure evaluates the Gram identity and the direct sum in the browser's native float64, the exact value with BigInt, and both kernels' enclosure radii as the package computes them. At the repo's frame (M = 1e6, delta = 1e-6) the Gram score is 0.0, the direct and exact scores are 1.0000152290447206e-12, the Gram enclosures overlap (hatched, undetermined) and the direct enclosures are disjoint (green, determined), so the verdict is REFUSED (GRAM_CANCELLATION). Switching to float32 collapses the two points into one stored vector, where 0.0 is the correct answer. Raising the precision is not the remedy; changing the formula is. The repo also measured torch.cdist at this frame (torch 2.14.0+cpu, quoted, not computed here): 0.0 with the matrix-multiply form and 1.00000761449337e-06 (a distance) with the direct form. Colour key: proof: enclosures disjoint: the pair is determined; refused: enclosures overlap, undetermined; ink: scores and interval outlines. The identity is exact in real arithmetic and badly behaved in floating point, because the last subtraction cancels when $x$ and $y$ are close to each other and far from the origin. The example I keep coming back to is two float64 points, $x = (10^6, 0)$ and $y = (10^6 + 10^{-6}, 0)$. The Gram identity returns exactly `0.0`. The direct sum $\sum_l (x_l - y_l)^2$ returns `1.0000152290447206e-12`, and so does exact scaled-integer arithmetic on the same stored bytes. separatrix, given the Gram scores, computes a radius of $1.776357 \times 10^{-3}$ around each score at $d = 2$, sees that the two enclosures overlap, checks that the direct kernel separates the pair, and returns `REFUSED (GRAM_CANCELLATION)`. At float32 the same two points are a single stored vector, since `np.float32(1e6 + 1e-6)` equals `np.float32(1e6)`, so `0.0` is the correct squared distance for the data as it stands and there is nothing to refuse. The frame is run in float64 precisely because the inputs are distinct there and the Gram identity still returns zero. The repository puts the lesson in one line: "Changing the formula is the fix; changing the precision is not." It matters what `CERTIFIED` does and does not mean. It means rounding did not choose this ranking. It says nothing about the model, the quantiser or the sensor that produced the vectors: a 384-dimensional embedding out of a float16 forward pass carries about $10^{-3}$ relative error, three to four orders of magnitude above the float32 rounding being certified. The certificate is about one named formula evaluated on the bytes it was handed, and nothing else. ## What it can do The headline is a statement about determinism across implementations. I took five corpora, 300 queries each at $k = 10$, and evaluated the same squared-distance formula on the same stored bytes nine different ways: numpy's Gram identity in float32, the same with the reduction order permuted, numpy's direct sum, numpy's Gram identity in float64, `scipy.spatial.distance.cdist`, `sklearn.metrics.pairwise.euclidean_distances`, `torch.cdist` on both of its compute modes, and `torch.cdist` at batch 32. None of those nine is separatrix's own arithmetic. **Measured: 0 of 1,116** certified top-10 sets that moved between any two of nine engines. Control: nine numerically distinct evaluations of one formula on the same stored bytes; all 8 sets that did move had been refused first. Interval: 1,500 decisions: 5 corpora × 300 queries, k = 10; 384 refused. n = 1,500. Source: 18c2eaf · WIN-16QAL06O9GB · CPython 3.11.9, numpy 2.4.6, scipy 1.17.1, torch 2.14.0+cpu · one bench.py run, seed 11. Of the 1,500 decisions, 1,116 were certified and none of those sets changed between any two engines. Eight sets did change, all on the synthetic clustered corpus, and all eight had been refused before the comparison was made. I want to be plain about the shape of that result. Four of the five corpora returned 0 in the "moved, refused first" column, which means that on those four the package had nothing to catch; only the clustered arm shows that the corpus was not too easy. Three of the five corpora are generated and two are downloads of a few thousand rows (MNIST and BEIR SciFact embedded with all-MiniLM-L6-v2). Every number in this block comes from one draw, seed 11, on one Windows machine. On a later build the benchmark was re-run and compared field by field: six fields differed, all six wall-clock timings, and every refusal count, agreement count and soundness field reproduced. The demo I find most persuasive is smaller. On a corpus of 400 float32 vectors in 64 dimensions, seeded with 40 near-duplicate pairs at a separation of $3 \times 10^{-7}$, separatrix looks at one evaluation and names the rows whose top-5 set it cannot decide. Then a second evaluation runs, the same formula with the columns permuted identically, which leaves every exact score unchanged and changes every rounding: ```text corpus 400 x 64 float32, 60 queries, k = 5, 40 near-duplicate pairs at 3e-07 named undetermined, from one evaluation: 9 of 60 [5, 13, 14, 22, 37, 39, 40, 54, 57] actually differed, over two evaluations: 3 of 60 [14, 22, 54] differed and NOT named: 0 [] ``` The last line is the only one that is evidence. A row that two evaluations decided differently and that was not named in advance would be a counterexample to the certificate, and the demo exits with code 2 on one. The control is in the corpus: a corpus made of nothing but near-duplicates would refuse every row and make the claim true by construction, so this one leaves 51 of 60 rows certified, each of which had something to lose. The 6 rows named but not moved are the bound's pessimism, printed in the same table as the 3 that were real. The one result scored against an answer key I did not compute is SIFT1M. It comes from a separate command, `bench.py --sift`, which downloads a 516 MB corpus once: 1,000,000 SIFT descriptors in 128 dimensions, 1,000 queries, $k = 10$, run with `chunk=100` in 48.0 s. The ANN_SIFT1M release ships its authors' exact top-100 lists, so the certificate can be checked against a third party. **Measured: 948 of 948** certified SIFT1M top-10 sets identical to the published ground truth. Control: the ANN_SIFT1M authors' ground-truth file, computed outside this repository. Interval: 52 of 1,000 refused (5.2%); the 9 rows where the published answer differs were all refused first, and all 9 are exact ties. n = 1,000 queries × 1,000,000 base rows. Source: 18c2eaf · separate run of bench.py --sift, chunk=100, 48.0 s · WIN-16QAL06O9GB, CPython 3.11.9. The nine rows where the published list and this float32 run disagree (93, 170, 460, 574, 614, 731, 760, 930 and 934) were all refused before the comparison, and exact arithmetic calls every one of them a tie. Row 93 is small enough to print: items `#196106` and `#274922` are both at squared distance 42,192 from the query in exact integers. The published list takes `#274922`, this run takes `#196106`, and the frontier reported `gap 0.000000e+00` before either was preferred. On a tie both answers are correct, and the right thing for a certificate to do there is refuse. (Note: The repository's `assets/data.json` records this run as `chunk` 20 and 145.43 s; that record comes from `make_assets.py --measure`, whose default chunk is 20. README and RESULTS quote the `chunk=100` run at 48.0 s. They are two runs, and I quote the second.) The clean 948 needs its caveat next to it, and the repository puts it there. SIFT descriptors are integers from 0 to 255; every Gram intermediate stays below $2^{24}$, which float32 carries exactly, so the float32 score is the exact integer distance. Measured, 0 of 20,000,000 float32 Gram scores differ from int64 arithmetic over the same bytes. The flip count on this corpus is therefore zero by arithmetic, and the table can test only whether the certificate ever contradicts a third party. It did not. A refusal is a return value, not an exception and not a crash. On the iid corpus at $d = 384$ the command-line tool prints this: ```text ============================================================================== REFUSED (BOUNDARY_UNDETERMINED) 285/300 determined ============================================================================== detail 15 of 300 rows have a rank-10 boundary this enclosure does not decide; the direct kernel separates 7 of those 15 frontier pairs computed kernel gram bound cheap per-row k 10 ------------------------------------------------------------------------------ boundary row 12: in #1046 [1.721866e+00, 1.722050e+00] out #1342 [1.721932e+00, 1.722116e+00] gap 6.616116e-05 width 1.840634e-04 deficit -1.179022e-04 14 further boundaries not shown (--max-report 1) ------------------------------------------------------------------------------ next The two enclosures at the rank-k boundary overlap. Pass escalate=True to decide it exactly, or bound='tight', or per_pair=True, or recompute in a wider dtype. ============================================================================== ``` There are four things to do with a refusal, and the README states what each costs. Exact escalation settles the boundary in scaled integers. Changing the kernel works when the refusal says the formula is the problem. Retrieving a few more results than needed lets whatever comes next absorb the boundary. And a genuine tie needs a tie-break rule that the caller owns, because no precision removes it. **Measured: 11 → 2** refused SIFT1M rows before and after escalate=True, 100 queries against the full base. Control: exact scaled-integer arithmetic on the frontier; 0 float sets moved. Interval: 22 exact products spent; the 2 left (rows 82 and 93) are exact ties, verdict NOT CERTIFIED (EXACT_TIE), exit 1. n = 100 queries × 1,000,000 rows. Source: 18c2eaf · bench.py --sift, 9.9 s · WIN-16QAL06O9GB, CPython 3.11.9. At $n = 10{,}000$ and $n = 100{,}000$ the SIFT refusals were `GRAM_CANCELLATION`, whose next action says to run the direct kernel. I ran it on the same 100 queries at $n = 100{,}000$: the Gram kernel refused 2 of 100 in 0.38 s, the direct kernel certified all 100 in 3.3 s. The advice works and it costs 8.7 times the time. The figure below runs the same mechanism on whatever you give it. Paste a small corpus or generate one, and it returns what `certified_topk` would: the status per query, the reason, the frontier pair with both intervals, and the next action. **Figure 2.** A port of certified_topk to the browser. The corpus and queries are yours (pasted) or generated: iid unit vectors, the clustered generator (20 points per cluster, spread 0.02), near-duplicate pairs, an integer lattice like SIFT, or sparse pixel values 0 to 255. Preconditions run in the package's order, then scores, radii, the rule per query and, if asked, exact BigInt escalation. Each query's ranked scores are drawn as intervals; the corridor between the largest member upper bound and the smallest non-member lower bound is drawn across the two rank-k boundary rows, green when open and hatched when it overlaps. float16 is emulated by rounding every operation to an 11-bit significand. The generators use the browser's own PRNG, not numpy's, so the live counts will not reproduce the repository's numbers. The refusal rate shown belongs to whatever was entered. Colour key: proof: CERTIFIED: corridor open between max-in and min-out; refused: REFUSED, enclosures overlap; ink: scores, intervals, members filled. The Python surface is four calls: ```python idx, v = separatrix.certified_topk(corpus, queries, k=10) # replaces argpartition j, v = separatrix.certified_argmin(corpus, queries) # k=1, axis squeezed e = separatrix.enclose_scores(corpus, queries) # the scores and their radii trit = separatrix.certified_threshold(e.D, e.R, t) # +1 / -1 / 0 undetermined ``` ## How it was made The construction has two layers that never see each other. The first produces, for every score, an interval guaranteed to contain the exact value. The second decides a set from those intervals alone, without knowing which kernel produced them. > **Definition: Enclosure.** > > Let $X$ (the corpus, $n \times d$) and $Q$ (the queries, $m \times d$) be the arrays as > stored in memory, and let $s$ be the value of the named score formula in exact real > arithmetic on those stored numbers. An enclosure is a pair $(D, R)$ of computed scores and > radii $R \ge 0$ with $|D_i - s_i| \le R_i$ for every $i$. The radius is an a-priori bound, > computed in the same pass as the scores: not a sample, not a second run, not a statistic. ### The bound Everything rests on one theorem from chapter 3 of Higham's *Accuracy and Stability of Numerical Algorithms*. With $\varepsilon$ the machine epsilon of the working dtype and $n$ the length of a reduction: $$ u = \frac{\varepsilon}{2}, \qquad \gamma_n = \frac{n\,u}{1 - n\,u}, \qquad \bigl|\,\mathrm{fl}(x^\top y) - x^\top y\,\bigr| \;\le\; \gamma_d \sum_{i=1}^{d} |x_i|\,|y_i| . $$ **Figure 3.** Tab A: for random float32 pairs, four legal reduction orders (sequential, reversed, pairwise, eight interleaved lanes) are evaluated with Math.fround after every operation, and each one's error against a float64 reference is plotted under the radius the package computes. Different orders give different numbers and the same radius; the readout counts escapes, which should be zero. This samples a class of evaluations and is weaker evidence than the theorem or the repository's exact-lattice gate. Tab B shows where the bound stops being a bound: the BOUND_VACUOUS wall where n·u exceeds 1/2, the RANGE_UNSAFE headroom check, the underflow term (repository numbers, quoted), the 4×4 canary, and the reduction-depth comparison behind the withdrawn pairwise model. The browser uses its own PRNG, so live counts will not reproduce the repository's. Colour key: ink: the radius R; parameter: actual error, one shade per reduction order; proof: inside the enclosure; refused: the bound refuses (vacuous or out of range). Two details of that statement carry weight. The absolute value sits inside the sum, on each product. The form $2|\langle x, y \rangle|$, which is tempting because it is cheap, is not the bound and falls below it exactly under cancellation; the module docstring records a 20.7× understatement at $d = 384$ on two random unit-norm vectors. And $\gamma_n$ is only meaningful while $n u$ is small. The code evaluates it in float64, rounds it outward with `nextafter`, and raises `BOUND_VACUOUS` when $n u > 1/2$, which the docstring describes as the point where the bound stops carrying information rather than where it inverts. For float16, RESULTS.md records $d = 1022$ as the last width that is not vacuous. The constant also covers every reduction order. Any reduction tree over $d$ products performs $d - 1$ additions, so its depth is at most $d - 1$, and $\gamma_{d+2}$ is an envelope over all of them. That is why batch size, thread count, chunking and FMA cannot escape it, and it is also why I deleted an earlier `summation="pairwise"` option (see Limitations). ### Two radii for the Gram identity The Gram identity is three reductions and two additions, plus an exact multiply by two, so its constant is $\gamma_{d+2}$ and its cross term is $2\sum_i |x_i y_i|$. Bounding that sum two ways gives two rungs: $$ R_{\text{cheap}} = \gamma_{d+2}\,\bigl(\lVert x \rVert + \lVert y \rVert\bigr)^2, \qquad R_{\text{tight}} = \gamma_{d+2}\,\bigl(\lVert x \rVert^2 + \lVert y \rVert^2 + 2\,\langle |x|, |y| \rangle\bigr). $$ **Measured: 0 of 656** enclosure escapes on the adversarial corpus. Control: exact squared distances in scaled Python integers (exact.exact_sq), no third-party arithmetic. Interval: 82 pairs in each of 8 configurations: 2 kernels × 2 bounds × 2 rungs, over 10 corpora. n = 656 pairs. Source: 18c2eaf · pytest tests/ then bench.py · WIN-16QAL06O9GB, CPython 3.11.9. By Cauchy–Schwarz, $\langle |x|, |y| \rangle \le \lVert x \rVert \lVert y \rVert$, so $R_{\text{cheap}} \ge R_{\text{tight}}$ pointwise and the ladder cannot invert; a self-check asserts it. The cheap rung needs only the two norms, which the identity already needs, so it costs one extra $O(nd)$ pass and no extra matrix multiply. The tight rung adds one matmul of absolute values. A memory-saving rung 1 replaces $\lVert y \rVert$ by $\max_j \lVert x_j \rVert$ for the whole query row, which is valid because $R$ is monotone in $\lVert x_j \rVert$. The norms are computed in float64. Reusing the working-dtype $\lVert x \rVert^2$ that already sits in the identity is the shortcut, and it is unsound in the direction that matters: a norm that rounds low gives a radius that is low, and a low radius is not a bound. Here is the code at the pinned commit (`separatrix/enclose.py:350-362`): ```python if bound == "cheap": if per_pair: R = g * (qn[:, None] + xn[None, :]) ** 2 else: R = g * (qn[:, None] + float(xn.max())) ** 2 else: absdot = np.abs(Q.astype(np.float64, copy=False)) @ np.abs( X.astype(np.float64, copy=False) ).T R = g * (qn2[:, None] + xn2[None, :] + 2.0 * absdot) if not per_pair: R = R.max(axis=1, keepdims=True) return _inflate(R, d, dt) ``` For unit-norm float32 vectors at $d = 384$ the cheap radius is $4\gamma_{386}$, with $\gamma_{386} \approx 386 \cdot 2^{-24} \approx 2.30 \times 10^{-5}$; the width of a pair is twice that, about $1.84 \times 10^{-4}$, which is the median Gram width the benchmark reports on the normalised corpora (1.841e-04). ### The direct kernel, and the asymmetry The direct sum $\sum_l (x_l - y_l)^2$ adds only non-negative terms, so there is no cancellation term at all and its bound is relative. Higham's theorem bounds the error in terms of the exact $s$; solving for the computed value gives a radius in terms of what is actually held: $$ R_{\text{direct}} \;=\; \max(D, 0)\,\frac{\gamma_{d+1}}{1 - \gamma_{d+1}} . $$ **Measured: 300 → 2** refusals on the clustered d = 384 corpus, Gram kernel against direct kernel. Control: the same 300 queries and the same stored bytes under the other kernel. Interval: 298 of the 300 Gram refusals attributable to the kernel alone; median width 1.841e-04 (Gram) against 9.167e-05 (direct). n = 300 queries. Source: 18c2eaf · bench.py seed 11 · WIN-16QAL06O9GB, CPython 3.11.9 · generated corpus. The Gram identity's radius is absolute and dominated by cancellation; the direct sum's is relative. The two kernels compute the same real number and have enclosures that differ by orders of magnitude on near pairs, and the module docstring calls that asymmetry "the most valuable thing this package can say". It is what lets a refusal name a code change. When the Gram kernel refuses a row and the direct kernel separates every refused frontier pair, the reason is `GRAM_CANCELLATION` and the next action is `kernel="direct"` or `torch.cdist`'s `donot_use_mm_for_euclid_dist` mode. The direct-kernel test sets the reason and never the certificate: certifying needs the full max-in/min-out comparison over all $n$, not a verdict on one pair. The code refuses when $\gamma_{d+1} \ge 1$ rather than returning a radius. That guard exists because of a measured bug, described in Limitations: at float16 and $d = 1023$, $(d+1)u$ is exactly $1/2$, $\gamma$ rounds outward to 1.0000000000000002, and dividing by $1 - \gamma$ produced a negative radius. ### What the radius's own arithmetic costs Two more terms ride on every radius. The relative-error model $\mathrm{fl}(ab) = ab(1+\delta)$ holds only for normal results; a product that lands subnormal carries absolute error. So an underflow term $\eta$ is added unconditionally. And the radius is itself computed in floating point (a float64 reduction, a square root, a sum, a square, a scale), so it is pushed outward by a derived factor. As the code computes them, with $s_{\min}$ the smallest subnormal of the working dtype and $u_{64} = 2^{-53}$: $$ \eta(d) = 4\,(d+2)\,s_{\min}, \qquad \mathrm{push}(d) = \gamma_{d+2}^{(64)} + 8\,u_{64}, \qquad R' = \mathrm{nextafter}\bigl(R\,(1 + \mathrm{push}(d)) + \eta(d),\; +\infty\bigr). $$ **Measured: 4000 → 0** trials where the exact score escaped the radius, float32 components near 1e-25, without and with η. Control: the ordinary regime, float32 components near 1: 0 of 4000 escaped with or without η. Interval: float16 components near 3e-4: 3078 of 4000 escaped without η, 0 of 4000 with it. n = 4,000 trials per regime, d = 8, seed 21. Source: 18c2eaf · pytest tests/test_enclose.py -k underflow · exact scaled-integer ground truth. The control row is what makes the underflow result a measurement rather than a patch: the escapes are specifically underflow, and $\eta$ is not covering for a broken bound in the ordinary regime. One inconsistency in the repository is worth naming here. The docstring of `eta` gives the rationale as "at most smallest_subnormal/2 per product", over the $d$ products. Summed, that would be $d\,s_{\min}/2$. The code's constant is $4(d+2)\,s_{\min}$, which is larger, and the factor $4(d+2)$ is not derived in the docstring. I have written the equation as the code computes it. The push factor has a derivation in its docstring, term by term, and a note that the obvious guess, $4u_{64}$, is short by about 200× at $d = 784$ ($4.44 \times 10^{-16}$ against $8.73 \times 10^{-14}$). The inflation is the last thing done to every radius (`separatrix/enclose.py:149-152`): ```python def _inflate(R: np.ndarray, d: int, work_dtype) -> np.ndarray: """R -> nextafter(R*(1+push) + eta, inf). The last thing done to every radius.""" out = R * (1.0 + _push(d)) + eta(d, work_dtype) return np.nextafter(out, np.inf) ``` ### Preconditions, before any score is read The bound covers evaluations that use the declared unit roundoff on finite, in-range inputs. So four preconditions run first, in a fixed order: P1, every input is finite (the refusal names the row and the column); P3, $\gamma$ is not vacuous; P2, nothing can overflow; P4, the multiplier really has the declared precision. Only then is the matrix multiplied. A precondition that runs after the scores, as the docstring puts it, certifies garbage: the float16 range case would read as a non-finite score, naming the damage instead of the cause. The range check is derived rather than a safety factor. By Cauchy–Schwarz every intermediate of the Gram identity, and every partial sum of the cross term, is dominated by one quantity per query row, computed in float64: $$ h_i \;=\; \bigl(\lVert q_i \rVert + \max_j \lVert x_j \rVert\bigr)^2 , \qquad h_i > \max(\text{working dtype}) \;\Longrightarrow\; \texttt{RANGE\_UNSAFE}. $$ **Measured: 100 of 100** SIFT1M query rows refused RANGE_UNSAFE at float16, before any score is read. Control: the float32 run on the same bytes, which certifies. Interval: max ‖x‖² 2.612e5 against float16's 6.55e4; the Gram intermediate (‖q‖ + ‖x‖)² reaches 1.04e6. n = 100 real queries. Source: 18c2eaf · bench.py --sift · WIN-16QAL06O9GB, CPython 3.11.9. Real MNIST shows the same thing at 5,000 rows, with maximum $\lVert x \rVert^2$ of $1.489 \times 10^{7}$. With `upcast=True` the package widens float16 to float32 and returns `CERTIFIED_UPCAST`, exit code 4 and never 0, because a build pipeline must not read a pass for a computation its index does not run. P4 is a 4×4 canary. Matrix $A$ holds $1 + \varepsilon$ everywhere; $B$ has two rows of ones and two of zeros, so every entry of $AB$ should be exactly $2(1 + \varepsilon)$, which is representable. Arithmetic that rounds its inputs coarser than the declared mantissa (TF32, bfloat16 inputs, Apple AMX) collapses $1 + \varepsilon$ to 1 and returns exactly 2, and the run is refused as `REDUCED_PRECISION_ARITHMETIC`. P5, the accumulator width, is not testable by either probe I tried, so it is a declared assumption that travels on every verdict as `accum_assumed`. ### The rule With an enclosure in hand, the decision is one inequality over $(D, R)$, knowing nothing about kernels. Let $T$ be the indices of the $k$ smallest $D_i$. Then $$ \max_{i \in T}\bigl(D_i + R_i\bigr) \;<\; \min_{j \notin T}\bigl(D_j - R_j\bigr) \;\Longrightarrow\; T = \mathrm{topk}(s), \qquad \begin{aligned} \text{gap} &= D_{\text{out}} - D_{\text{in}}, \\ \text{width} &= R_{\text{in}} + R_{\text{out}}, \\ \text{deficit} &= \text{gap} - \text{width} \;>\; 0 , \end{aligned} $$ **Figure 4.** An exact port of decide.py. Each score is a tick with its interval \[D - R, D + R]; members of T are filled. The corridor between the largest member upper bound (max-in) and the smallest non-member lower bound (min-out) is green when the rule holds and hatched when it does not. The grey bracket marks the pair a naive rule would compare, ranks k and k + 1. The preset \[0, 1, 2, 10] with radii \[12, 0, 0, 0] and k = 2 is the repository's counterexample: the naive pair certifies {0, 1}, the rule refuses with max-in 12.0 against min-out 2.0, and the repository's witness (11, 1, 2, 10) lies inside the box with top-2 {1, 2}; the diamonds mark the box's own worst corner (12, 1, 2, 10), which has the same top-2. In 3D mode (three scores) the box and the walls v_i = v_j show the same rule as geometry. The 2,000-random-box check uses the browser's own PRNG, so its count will not match the repository's. Colour key: proof: rule satisfied: corridor open; refused: intervals overlap, not determined; baseline: the naive rank-k / rank-(k+1) pair; ink: scores; members of T filled; diamonds mark the worst corner. Here "in" is the member of $T$ with the largest upper bound and "out" is the non-member with the smallest lower bound. Those two indices, and no others, are what the rule compares, and they are the pair a refusal names. The inequality is strict: equality is not determined. The implementation is short (`separatrix/decide.py:129-152`): ```python R = broadcast_radius(D, R) T = topk_set(D, k, largest=False) mask = np.zeros(D.shape, dtype=bool) mask[T] = True hi = D + R lo = D - R inside = int(T[np.argmax(hi[T])]) out_idx = np.flatnonzero(~mask) outside = int(out_idx[np.argmin(lo[out_idx])]) f = Frontier( row=row, inside=inside, outside=outside, inside_lo=float(lo[inside]), inside_hi=float(hi[inside]), outside_lo=float(lo[outside]), outside_hi=float(hi[outside]), gap=float(D[outside] - D[inside]), width=float(R[inside] + R[outside]), ) if not f.determined: return f ``` Why the rule implies determinism is a short argument. Every vector in the box $\prod_i [D_i - R_i, D_i + R_i]$ has its members of $T$ below max-in and its non-members above min-out, so when the two sides are disjoint every vector in the box has top-$k$ set $T$. The exact scores are in the box. So is every other evaluation of the same formula on the same stored bytes whose own error is covered by the bound. Batch size, BLAS backend, thread count, chunk size, reduction order, FMA and `torch.cdist`'s 25-row switch therefore cannot change $T$. The package docstring puts it as "Determinism across backends is a theorem here, not a measurement over reruns." Checking the rule on the worst corner of the box (members pushed up, non-members pushed down) is a proof over the whole box, not a sample of it. The rule uses the unclamped interval, even though a squared distance cannot be negative, because the rule does not know the kernel and an inner-product score can be negative. I measured what that costs: across all five benchmark corpora, 0 of 4,764,900 lower bounds fell below zero, so the clamp would change no refusal. `largest=True` negates the scores and reuses the same rule, so there is one rule and one place for it to be wrong. `ordered=True` additionally requires the $k - 1$ adjacent member intervals to be disjoint; adjacency suffices because "entirely to the left of" is transitive. ### Escalation to exact arithmetic A refusal can be settled exactly, but not by escalating only the pair it names: a third index's enclosure can still straddle. The repository's example is scores $[1.0, 1.05, 1.06, 5.0]$ with radii $[0.1, 0.1, 0.1, 0]$ and $k = 2$; resolving the named pair exactly still leaves index 0 straddling. So escalation works on the frontier $$ F = \bigl\{\, i \in T : D_i + R_i \ge \text{min}_{\text{out}} \,\bigr\} \;\cup\; \bigl\{\, j \notin T : D_j - R_j \le \text{max}_{\text{in}} \,\bigr\}. $$ **Measured: 9 of 11** refused SIFT1M rows decided by frontier escalation. Control: exact scaled-integer distances on the frontier indices. Interval: 22 exact products; the other 2 rows are exact ties; 0 float sets moved. n = 100 queries × 1,000,000 rows. Source: 18c2eaf · bench.py --sift, escalate=True · WIN-16QAL06O9GB, CPython 3.11.9. Escalation computes exact squared distances for every index in $F$, re-decides, and adds new frontier indices until the set closes, the two extreme scores are exactly equal, or a budget (`max_escalations`, default 64) runs out. The frontier is computed in floats with a one-ulp cushion in the permissive direction, so it can only gain indices, never lose one. There are three outcomes: the set closes and is returned; the exact scores tie, which is reported as `NOT CERTIFIED (EXACT_TIE)`, exit code 1, the one outcome no precision removes; or the exact set differs from the float set, in which case the corrected indices are returned with `float_set_differed` set, because returning the float set there would contradict the guarantee. The exact scores avoid `fractions.Fraction`, which runs a gcd on every addition. Every float is $m \cdot 2^{e}$, so scaling by a fixed power of two makes it an integer with one shift: $$ x \;\mapsto\; \mathrm{int}\bigl(x \cdot 2^{b}\bigr), \qquad b = 149 \ (\text{float32, float16}), \quad b = 1074 \ (\text{float64}), \qquad \lVert x - y \rVert^2 \ \text{returns scaled by } 2^{2b}. $$ **Measured: 0** escalations that contradict a certificate. Control: exact_sq re-decision of every refused row. Interval: all 384 refused rows of the benchmark, 5,827 exact dot products. n = 384 rows. Source: 18c2eaf · bench.py seed 11 · WIN-16QAL06O9GB, CPython 3.11.9. The same scaled-integer code is both the benchmark's ground truth and the escalation a user runs, so the two cannot drift apart. The statuses map to exit codes: 0 `CERTIFIED`, 1 `NOT CERTIFIED`, 2 `REFUSED`, 3 a usage error, 4 `CERTIFIED_UPCAST`. A `gate(max_refused, fixture)` context lets a build fail on the refused fraction; `max_refused` has no default, and on a mismatch of its nine-field configuration digest the gate refuses to compare rather than going red. ## What's new in it The usual way to check this kind of output is to run it twice: once in float32 and once in float64, and diff the lists. I measured that practice beside the package, and the honest first sentence is that it is mostly fine. It was right on 1,495 of 1,500 rankings. What it cannot do is say anything about a row where both runs happened to agree. The diff needs two runs and proves nothing when it comes back empty. separatrix names, from one evaluation, every row that rounding could change, and proves the rest. **Figure 5.** Panel A replays the repository's frame 3 in the browser: a 400 × 64 float32 corpus with near-duplicate pairs, 60 queries, k = 5. From evaluation A alone, each query cell is either certified (green) or named undetermined (hatched). Then further evaluations run (columns permuted, reversed order, pairwise tree, eight lanes, all emulated with Math.fround), and a dot marks every row whose top-k set moved. A dot on a green cell would be a counterexample and is drawn as an X. The corpus uses the browser's own PRNG, so the live counts will not reproduce the repository's 9 named, 3 moved, 0 not named, which are printed beside the live run. Panel B is the nine-engine table from the repository, verbatim; only the clustered arm carried evidence. Colour key: proof: certified from one evaluation; refused: named undetermined; ink: dot: set moved between evaluations; X: counterexample; parameter: evaluation chips A to E (gold): the engines that ran. The second usual approach is to make runs agree. `torch.use_deterministic_algorithms`, `CUBLAS_WORKSPACE_CONFIG` and reproducible BLAS libraries give bitwise run-to-run reproducibility. They are free or cheap and they are in-tree. They also make two runs agree on a value that may never have been determined. separatrix does the opposite: it does not make the numbers agree, it lets them differ freely and proves that the set cannot. The third is to raise the precision. The cancellation frame above is the reason I do not trust that as a fix. In float64, two distinct points still come back at exactly zero under the Gram identity. A refusal coded `GRAM_CANCELLATION` names the formula as the problem, and the measured fix (the direct kernel, 8.7× the time on SIFT) is a change of formula. The fourth is the obvious interval rule, and this is the part I found most instructive to build. Once scores have intervals, the natural check is to compare only the rank-$k$ and rank-$(k+1)$ intervals: $$ D_{(k)} + R_{(k)} \;<\; D_{(k+1)} - R_{(k+1)} \quad (\text{naive, unsound when radii vary}). $$ **Figure 6.** Five panels of 2,000 random boxes each (n = 12, k = 3, scores standard normal), with radii R_i = c · U_i^p for p = 0 to 4. At p = 0 the radii are constant and the naive pair rule and the sound rule agree on every box. As p grows the radii become unequal and the naive rule starts certifying boxes the sound rule refuses; the figure searches each such box for a witness vector inside it with a different top-k set. The sixth panel is the repository's result on real benchmark corpora: identical refusal counts under both rules. All live counts come from the browser's own PRNG and are derived, not repository numbers. Colour key: proof: certified by both rules; baseline: certified by the naive rule only; ringed if refuted by a witness; ink: radii histogram per panel; repository refusal counts; refused: refused by both. That rule is a false theorem. The radius scales as $(\lVert q \rVert + \lVert x \rVert)^2$, which depends on the norm of each candidate and has nothing to do with the order of the scores, so radii vary across a row. With scores $[0, 1, 2, 10]$, radii $[12, 0, 0, 0]$ and $k = 2$, the naive pair is indices 1 and 2: $1 < 2$, disjoint, certified $\{0, 1\}$. But the vector $(11, 1, 2, 10)$ lies inside the box and its two smallest are $\{1, 2\}$. The sound rule compares max-in $12.0$ against min-out $2.0$ and refuses. What makes this worth an essay section is that on every benchmark corpus I ran, the naive rule and the sound rule returned identical refusal counts: 15/15, 300/300, 35/35, 32/32 and 2/2. The naive rule would have shipped green. Only the counterexample test catches it, which is why that test exists and why the naive rule is shipped in `decide.py`, never called by anything that certifies, so the test scores against real code rather than a retyped formula. The fifth is the a-posteriori bound. Ogita, Rump and Oishi's Dot2 is far tighter than any $\gamma_n$, but it is elementwise and forfeits the matrix multiply, while $\gamma_n$ needs only norms and rides the one BLAS call the scores already cost. That is a throughput decision; I ran Dot2 as a benchmark arm expecting it to win on coverage and lose on throughput, and it did both. Finally, a refusal here is a typed return value with eight codes, each naming a concrete object: a boundary pair, a row and column, a dtype, a budget. There is no `on_refuse` handler and no default threshold, because, as the README says, a library that throws on 26% of rows is uninstalled the same afternoon. ## What no one else built The README opens its prior-art section with this: "Nothing in the certified layer is new, and saying so first is what makes the rest believable." The bound is textbook, the rule's shape is old in computational geometry, and the README counts three prior works that beat this design on their own axes. What I can claim is a specific combination aimed at a specific adversary. **The bound itself.** Higham's [*Accuracy and Stability of Numerical Algorithms*](https://doi.org/10.1137/1.9780898718027), chapter 3, gives Theorem 3.1 and the $\gamma_n$ notation. separatrix adds no new inequality to it. What it adds is the plumbing from that bound to a set decision with a typed refusal, plus the subnormal term $\eta$ and the derived push on the radius's own rounding. **Verified interval computing.** Rump's [INTLAB](https://www.tuhh.de/ti3/rump/intlab/) is a MATLAB and Octave toolbox for verified computation with real and complex intervals, vectors and matrices. It carries intervals through its arithmetic. separatrix does not carry intervals through arithmetic: it computes ordinary float scores with the same BLAS call the caller would use, attaches one a-priori radius per score from the norms, and then decides a set. That is far narrower than INTLAB, and the narrowness is what lets the whole check ride one BLAS call. **Accurate dot products.** Ogita, Rump and Oishi, ["Accurate Sum and Dot Product"](https://doi.org/10.1137/030601818) (SIAM J. Sci. Comput., 2005\), give Dot2 and an a-posteriori bound near $u$ where separatrix's a-priori bound is near $(d+1)u$. Benchmarked on 64 clustered queries in one draw, Dot2 refused 0 of 64 boundaries where separatrix refused 64 of 64, at 61× to 99× the cost over six runs. Anyone who can afford roughly 100× should use Dot2. The difference is throughput only. **Filtered predicates.** CGAL's [`Filtered_predicate` and `Interval_nt`](https://doc.cgal.org/latest/Number_types/index.html) have shipped this exact pattern since around 2001: interval enclosure, disjoint certifies the sign, overlap escalates to exact. Melquiond and Pion [formally certified the static filters](https://doi.org/10.1051/ita:2007005), the same class of a-priori bound, machine-checked, where mine are only tested against an integer oracle. The concrete difference is the predicate. A geometric filter decides the sign of one expression in 2 or 3 dimensions. separatrix decides a rank-$k$ set boundary over $n$ candidates in hundreds of dimensions, which needs max-in/min-out over the whole set rather than one pair, and needs escalation over a frontier rather than over one expression. **Reproducible arithmetic.** [ReproBLAS](https://bebop.cs.berkeley.edu/reproblas/) (Demmel, Ahrens and Nguyen) produces bitwise-identical sums and BLAS results independent of summation order and processor count; `torch.use_deterministic_algorithms` does the same within one library. Both make runs agree. Neither reports whether the answer was ever in doubt, and neither covers a different library on the same bytes. separatrix leaves the numbers free to differ and proves the set invariant over every evaluation inside the bound, across libraries; the nine-engine table is that claim measured. **Stochastic arithmetic and probabilistic bounds.** [CADNA](http://cadna.lip6.fr/) counts unstable branches with discrete stochastic arithmetic, which is the same output shape and older; its answer is a confidence estimate from a few random-rounding runs, where separatrix's is a proof over a box. Higham and Mary's [probabilistic rounding error analysis](https://doi.org/10.1137/18M1226312) would cut the width by about 19× at $d = 384$ and would very plausibly close much of the pessimism gap. I cite it and do not use it, because it is a probability and no output of this package carries one. **Certified top-k in machine learning.** Jia, Cao, Wang and Gong, [Certified Robustness for Top-k Predictions against Adversarial Perturbations via Randomized Smoothing](https://iclr.cc/virtual/2020/poster/1946) (ICLR 2020), certify that a classifier's top-$k$ labels are stable under bounded input perturbations. The name is close and the object is different: their adversary is an attacker moving the input and their guarantee comes from Monte Carlo sampling of a smoothed classifier, while separatrix's adversary is rounding on fixed stored bytes and its guarantee is deterministic. scikit-learn's [PR 13554](https://github.com/scikit-learn/scikit-learn/pull/13554) (chunked upcasting in `euclidean_distances`, on by default) is the mitigation for the diagnosis this package makes, and it shrinks the population that needs separatrix. Against that list, the pieces I built and did not find in the work I compared are these. A set rule with a pinned proof that the obvious pair rule is unsound when radii vary, and a test that catches it where every benchmark could not. Escalation over the frontier set $F$, with a third outcome that returns the exact set and flags that the float set was wrong. The kernel asymmetry turned into a refusal code that names a change to the formula. Preconditions, including a 4×4 canary for reduced-precision multipliers, that refuse before any score is read. And a refusal that is a value: the boundary pair, both intervals, the gap, the width and the deficit, from one evaluation. I make no claim about whether someone has built this combination elsewhere; I claim only the differences above, against the works named. **Figure 7.** The same near-duplicate corpus and the same float32 scores, decided four ways, each scored against exact BigInt arithmetic: separatrix's rule, the float32-vs-float64 diff, a tuned margin gap > eps · |score|, and the naive pair rule. Each column is a query. Green is certified and right, grey ringed is certified and wrong against exact arithmetic, hatched is refused or unflagged, and an ink dot is a row the diff flagged. The eps slider and the small multiples sweep the tuned margin. Green bars are certified and right; the dashed line is separatrix's certified count. Rows with an exact tie have no unique truth and are reported as ties. This is one constructed corpus from the browser's own PRNG; its counts will not reproduce the repository's and must not be quoted as repository results. Colour key: proof: certified and right; baseline: old practice; ringed when certified but wrong; refused: refused, or no flag; ink: dot: flagged by the diff; ink: dashed: separatrix certified count (eps multiples); ink-3: outlined t: exact tie at the boundary, no unique truth. The one measurement that tests the certificate against an answer neither I nor my code produced is SIFT1M, and the figure below lays it out row by row. **Figure 8.** SIFT1M, 1,000 queries against 1,000,000 vectors, k = 10, from the repository's separate bench.py --sift run (data in assets/data.json at 18c2eaf). Panel A is the empirical distribution of the rank-10 gap against the enclosure width (median 16.12): 52 rows fall at or below the width and are refused, 948 clear it and are certified, and all 948 match the published top-10. The 9 rows where the published answer differs are ringed; all have gap 0 and all were refused first. Panel B is the rank ladder for query 0 (gap 1,483 against width 16.11, certified) and query 93 (items 196106 and 274922 both at 42,192, refused). Panel C shows what scale did: refusals at n = 10k, 100k and 1M, memory before and after the chunking fix, and the cost of the direct-kernel advice. The 948 is clean because SIFT's integer components make float32 scores exact (0 of 20,000,000 differ from int64), so this corpus cannot show rounding choosing a ranking. Colour key: proof: certified: gap clears the width; refused: refused, gap within the width; ink: ECDF, intervals; rings mark the 9 disputed rows; measured: panel C: repository-measured scale, memory and time. On the 100-query block the refused count grew with $n$: 2 at 10,000 rows, 2 at 100,000 and 11 at 1,000,000. That growth was the prediction on record before the run, because a fixed-width enclosure has more chances to straddle the rank-$k$ boundary as the corpus fills in around it. It is a statement about the corpus, not the bound. At a million rows the median refused frontier had a gap of 7.0 against a width of 16.12, with a true error of 0. ## Limitations These are stated as the repository states them, collected here once. The certificate covers the rounding of one named formula on the bytes it was handed, and is blind to model, quantiser and sensor error (see What it is). It is not a detection. The float32-against-float64 diff was right on 1,495 of 1,500 rankings measured. The pre-registered prediction that no corpus would flip failed on the clustered arm, where 5 of 300 sets flipped, and the diff finds the same 5. On that arm both saw the same thing and only one of them needed a second run. It is not faster than what it replaces. The per-row rung costs 0.0529 s on the printed draw against 0.0314 s for the float64 Gram control and 0.0498 s for the diff: 1.68× the control and 1.06× the diff. Over six runs the ratios were 1.59× to 1.89× the control and 0.96× to 1.12× the diff. The honest words are roughly 1.7× the float64 control and parity with the diff. It is pessimistic where it matters. Of the 1,500 benchmark decisions, 384 were refused (26%), 300 of them on the synthetic clustered corpus, and escalation decided that 379 of the 384 were pessimism rather than damage; the other 5 are the clustered flips. On the exact integer lattice the smallest margin certified is 1,024 and the largest margin the float decision got wrong over 8 trials is 4, a factor of 256×, which is an upper bound on the gap and not a measurement of it. The refusal actions that work offline, exact escalation and a tie-break, do not ship at that refusal rate; the action the README recommends for production, without measuring it, is to retrieve $k + 5$ and let the reranker absorb the boundary. **Figure 9.** Small multiples at d = 16, 64, 256 and 1024 (300 corpus rows, 60 queries, k = 10, float32 emulated). Each panel sorts the queries by gap over width, with the line at 1 where the deficit is zero, for the Gram kernel per row, the Gram kernel per pair and the direct kernel. A norm-spread slider makes row norms vary, which is where the per-row collapse costs refusals; float16 shows the BOUND_VACUOUS wall near d = 1022. The corpora are synthetic Gaussian rows from the browser's own PRNG, so the live refusal counts will not reproduce the repository's; the repository's anchors at d = 384 and 784 can be overlaid as ghost markers. A refused strip is not a flip, and refusal rate is not a quality score. In the lower plot the solid line is the Gram kernel per row and the dashed line the Gram kernel per pair. Colour key: proof: determined: gap over width above 1; refused: refused: hatch (Gram), dashed outline (direct); ink: filled ticks Gram, outlined ticks direct. Two competitors beat it on coverage. Dot2 refused 0 of 64 boundaries where separatrix refused 64 of 64, at 61× to 99× the cost; the denominator is a 3 ms gemm whose noise dominates the ratio, and only the direction is not in doubt. A four-line tuned margin, certify when $\text{gap} > \varepsilon\,|\text{score}|$, certifies 295 of 300 on the clustered arm where separatrix certifies 0. What it lacks is any $\varepsilon$ below $10^{-3}$ that avoids certificates a witness contradicts, and the $\varepsilon$ that does costs 122 of 300 certificates on the iid arm and 108 on MNIST-shaped. The repository states the differentiator at its real size: "no tuning, and a proof." **Figure 10.** The arms that lost, at the same size as the ones that won, all verbatim from RESULTS.md at 18c2eaf. Tab 1: refusals by corpus under the Gram and direct kernels, with the totals 384 of 1,500 refused and 379 of 384 refusals pessimism. Tab 2: the tuned margin gap > eps · |score| across eps = 1e-8 to 1e-3 on five corpora, certified count and witnessed-wrong count; no eps below 1e-3 avoids wrong certificates on all five. Tab 3: cost, best of 5 on a generated 5,000 × 784 float32 corpus, and the Dot2 arm (0 of 64 refused against 64 of 64, at 61× to 99× the cost), with struck timings left out of the bars. Tab 4: torch.cdist's two compute modes are bit-identical at 24 and 25 rows and differ by 9.766e-04 at 26. Colour key: baseline: the competing practice: diff, tuned margin, Dot2; parameter: time: seconds and cost ratios; proof: separatrix where it certifies; ink: refusal counts, widths and errors. Several results are synthetic only: the nine-engine table apart from its two downloaded corpora, the tuned-margin baseline, the Dot2 arm, the shuffled-enclosure control, the cost table, the pessimism factor, and the chunk effect. The synthetic clustered generator (20 points per cluster at spread 0.02) is far more adversarial than a real index; it refused 300 of 300 where real BEIR SciFact refused 2 of 300, and no number from it should be read as a statement about a real index. The one real corpus at scale, SIFT1M, is arithmetically easy, so its flip count is 0 by arithmetic and it can never produce evidence that rounding chose a ranking. The nine-engine comparison was not run at a million rows, because nine evaluations of a 1,000 × 1,000,000 score array is 72 GB. The memory-saving per-row rung costs refusals where norms vary: 35 against 26 on MNIST-shaped and 32 against 15 on real MNIST, and 0 at $d = 384$ with unit norms. `ordered=True` compares $k$ boundaries instead of one and refuses more often. The argmin and threshold surfaces ship, but neither has a refusal count on any corpus. What is not claimed anywhere: a GPU number (there is no CUDA path; a CUDA tensor forces a host copy before the enclosure); any approximate index (IVF, PQ and HNSW search error is about $10^{-2}$ relative, four orders above the float32 enclosure width, so certifying their rounding would certify the wrong quantity); the accumulator width, which is assumed; a machine-checked radius, which Melquiond and Pion have for static filters and I do not; and any probability. Everything was measured on CPython 3.11.9 on one Windows machine. The repository states a suite of 212 tests, 212 passing with the SIFT1M cache present and 211 passing with 1 skipped without it; I report that count as the repository states it and have not reconciled it independently. **What failed.** - Withdrawn: separatrix costs 0.14× the fp64 reference (and the revision, 0.91× fp64 / 0.54× the diff) Killed by: did not reproduce; measured 1.68× the float64 control and 1.06× the diff, 1.59×–1.89× over six runs (RESULTS.md §6). - Withdrawn: summation="pairwise" as a certificate Killed by: OpenBLAS reduction depth is about d/8 + 3 ≈ 101 at d = 784 against the pairwise model’s 21.2: 4.8× too shallow by counting; option deleted. - Withdrawn: flipped = 0 on every corpus (pre-registered) Killed by: the clustered arm flipped 5 of 300 (rows 91, 96, 212, 236, 274), and the float32-vs-float64 diff finds the same 5. - Withdrawn: 107/300 refused on a clustered d = 384 index, generalising to real embeddings Killed by: 300/300 on the synthetic clustered generator and 2/300 on real BEIR SciFact. - Withdrawn: push = 4u₆₄ covers the radius’s own rounding Killed by: short by about 200× at d = 784 (4.44e-16 against 8.73e-14); replaced by γ\_{d+2}(f64) + 8u₆₄. - Withdrawn: a Verdict may cap its frontiers at 64 Killed by: 5 refused rows past the cap read as certified (5 apparent disagreements before, 0 after); the cap was removed. - Withdrawn: the direct kernel’s relative radius is sound wherever P3 admits n·u ≤ 1/2 Killed by: at float16, d = 1023, (d+1)u = 1/2 gave radius −9.224903e+14 and certified every row; now BOUND_VACUOUS. - Withdrawn: rung-1 timings of 0.0650 s and 0.0505 s; Dot2 ratios of 48×–78× and 81×–117× Killed by: none reproduced on the current build; struck, and the ratio bands kept as the measurement. SIFT1M also broke three things, none of them soundness: escalated frontiers all reported row 0, `chunk=` was validated and then ignored (an 8.00 GB allocation with 2.9 GB free), and the norm pass copied the corpus (1.02 GB). After the fixes the peak working set was 2,565 MB. The repository also disagrees with itself in places. The `decide.py` docstring quotes older prototype counts (19/19, 107/107, 20/20) and an unmeasured 40% ordered refusal rate; the `eta` docstring does not derive the code's constant; and the README's roadmap calls frontier escalation "the missing piece" although `escalate=True` already implements it. ## Read more [View the project](https://github.com/teerthsharma/separatrix) · [Source on GitHub](https://github.com/teerthsharma/separatrix) The source is at [github.com/teerthsharma/separatrix](https://github.com/teerthsharma/separatrix). Short link: [teerth.dev/separatrix](https://teerth.dev/separatrix). Every line number in this essay is at commit `18c2eaf`. The files worth reading, in order: - [README.md](https://github.com/teerthsharma/separatrix/blob/18c2eaf31f7a6ecc8cca44d54a3f884d67bc5c43/README.md): the plain-words version, the exact statement, the prior-art table and the limits. - [RESULTS.md](https://github.com/teerthsharma/separatrix/blob/18c2eaf31f7a6ecc8cca44d54a3f884d67bc5c43/RESULTS.md): every table with its command and control, the arms that lost, and the SIFT1M section. - [separatrix/decide.py](https://github.com/teerthsharma/separatrix/blob/18c2eaf31f7a6ecc8cca44d54a3f884d67bc5c43/separatrix/decide.py): the rule over $(D, R)$, the false theorem and the worst corner. - [separatrix/enclose.py](https://github.com/teerthsharma/separatrix/blob/18c2eaf31f7a6ecc8cca44d54a3f884d67bc5c43/separatrix/enclose.py): $\gamma_n$, $\eta$, the push, the preconditions and both kernels' radii; the only file whose bugs are unsound rather than merely wrong. - [separatrix/exact.py](https://github.com/teerthsharma/separatrix/blob/18c2eaf31f7a6ecc8cca44d54a3f884d67bc5c43/separatrix/exact.py): the scaled-integer oracle and frontier escalation. - [bench.py](https://github.com/teerthsharma/separatrix/blob/18c2eaf31f7a6ecc8cca44d54a3f884d67bc5c43/bench.py): the benchmark, including the Dot2 and tuned-margin arms. Related essays on this site: - [Epsilon-Hollow](/epsilon-hollow), the project separatrix came out of. - [tangle](/tangle) and [planimeter](/planimeter), two other tools of mine that certify an answer or refuse to give one. --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # planimeter > An agent edits a drawing and reports the room count. planimeter reads the file and returns the count, or the coordinate where the file cannot say. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/planimeter - Repository: https://github.com/teerthsharma/planimeter - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "planimeter", 2026, https://teerth.blog/planimeter - Languages: Python - Question: Did the number of enclosed rooms in this drawing change, and if the file cannot settle it, where exactly is the doubt? - Headline result: 0 wrong integers on 528 constructed drawings: 495 exact, 33 refused (control: same 528 files: round-6 snap 350 wrong, shapely polygonize_full 336 wrong) ## What it is An agent edits a drawing and reports the edit. "I added a wall that splits the room in two." The person who asked wants to know one thing: did the number of enclosed rooms change? planimeter answers that question for any SVG line drawing an agent writes. It reads the lines out of the file, never draws a picture, and returns three integers: **pieces**, the number of connected parts of the line work; **faces**, the number of enclosed regions; and **χ**, the Euler characteristic, which for planar line work is pieces minus faces. Getting those three integers from a correct graph is easy. Euler's identity gives all three from the vertex count, the edge count and the number of connected pieces, and the code that does it is one union-find pass. The README says it plainly: "This is 18th-century arithmetic and it is not where the risk is." The risk sits one step earlier. A drawing is a list of segments with floating-point endpoints, and two endpoints written $10^{-7}$ apart may be one corner drawn twice or two corners that happen to be close. Merge them and a wall closes a room. Keep them apart and the wall is a dangle. So the hard part, in the README's words, is "deciding which endpoints are the same endpoint". planimeter makes that decision from the data. It lists every separation the answer could turn on (corner to corner, and corner to a wall it does not touch), sorts them, and looks for a wide empty stretch: two consecutive values whose ratio is at least $\rho = 10$. Below that stretch is the drawing's wobble. Above it is real geometry. Endpoints closer than the bottom of the stretch are one vertex. The grouping cannot change anywhere inside the stretch, so it survives a tenfold change in tolerance, and the tool prints the stretch on every answer. When there is no such stretch, or when a line end sits inside the doubtful range beside a wall it does not quite touch, planimeter does not pick one of the two integers. It refuses. The refusal gives the reason, the coordinate in the file's own units, the id of the element that owns it, and one action. A refusal is falsy in Python, and asking for its integer raises `TypeError`. The README puts the rule in five words: "A refusal is never a number." The case that motivates all of this is small. Take a 10-unit square and a crosswall from the top edge down to a point $2.2\times10^{-6}$ above the floor. If the gap is closed, the wall divides the room and there are 2 faces. If it is left open, the wall is a dangle and there is 1 face. "Nothing in the file says which." A rounding snap, a GEOS polygonizer or a rasteriser will each return one of the two integers, because returning an integer is what they do. planimeter returns the coordinate `(5, 2.2e-06)` and the band `(0, 5)` the end sits in. **Figure 1.** A crosswall ends a distance g above the floor of a 10-unit square, swept here on a log axis. The plot shows what planimeter reports at each g: a certified integer, or a refusal with the coordinate it turns on (hatched band). Beside it is the round-6 strawman's integer. The two dashed lines are the two possible readings, closed (2 faces) and open (1 face). Only g = 0 and g = 2.2e-6 are repo fixtures. The other gaps are this figure's own sweep, run on a port of the package, and the readout at 2.2e-6 should match the README's refusal block. An inset magnifier shows the wall end; a button jumps to g = 0. Colour key: proof: certified integer from planimeter; baseline: round-to-six-decimals snap (the benchmark labels it STRAWMAN); parameter: the gap g you move; refused: refused: the tool names the coordinate and returns no integer. The intended user is an agent, not a person at a CAD station. `planimeter init` installs the tool as a `PostToolUse` hook in the agent's harness. After every write to an `.svg` file, the hook stamps one line into the agent's context, for example `planimeter walls.svg pieces 1 faces 4 -> 5`, and stays silent on every other write. The agent never has to look at the picture to know what it drew. The README credits my project [tangle](/tangle) as the template for the shape of the tool: an exact integer, a one-directional certificate, a typed refusal that names what to re-observe, and a benchmark with controls. ## What it can do The headline is on a constructed stratum, where the truth is known because the drawings were built. There are 44 figure families: segments, triangles, squares, a T-junction, an H, a noded pentagram, trees, grids from 2×2 to 12×12, and theta, chain, comb and ladder graphs. Each family is jittered at six levels of $\sigma/g$ from $10^{-7}$ to $10^{-2}$, where $g$ is the figure's own smallest feature gap, with two seeds each. That gives 528 drawings. A test parses `corpus.py` and asserts that it imports nothing from the package's counting, snapping or arrangement code, so the answer key cannot borrow the arithmetic it grades. **Measured: 0 wrong** planimeter on the jitter stratum: 495 exact, 33 refused. Control: CONTROL refuse-on-anything: 0 wrong, 0 exact, 528 refused. The zero only means something next to the 495 exact.. n = 528 draws (44 families × 6 jitter levels × 2 seeds). Source: RESULTS.md §1, bench.py run of record at 5713012 on WIN-16QAL06O9GB (Windows 11, Python 3.11.9), as reported by the repo; not reproduced here. The same 528 files, through the arms the repo ran against it: **Measured: 350 wrong** round-to-six-decimals dict snap + union-find, labelled STRAWMAN by the benchmark: 178 exact, 0 refused. Control: planimeter on the same 528: 0 wrong. networkx's b1 after the same snap also gets 350, because it computes the same E − V + C on the same snapped graph.. n = 528. Source: RESULTS.md:61-64 @ 564427f; run of record 5713012 as reported. **Measured: 336 wrong** shapely polygonize_full(unary_union): 192 exact, 0 refused (shapely 2.1.2, GEOS 3.13.1). Control: set_precision(1e-6) first: 299 wrong. polygonize given planimeter's own certified radius (the oracle control): 22 wrong, 473 exact, 33 refused.. n = 528. Source: RESULTS.md:65-75 @ 564427f; run of record 5713012 as reported. Two controls make these rows readable. The refuse-on-anything control also scores zero wrong, which is why the README never prints planimeter's zero without its 495. The oracle control was built to show that the counting layer adds nothing. It hands shapely planimeter's own certified radius and lets GEOS do the rest. If it had scored zero, the whole contribution would have been the printed tolerance. It scored 22 wrong. RESULTS.md describes them as `chain(k)` and similar figures, where the shared-edge structure survives snapping but not GEOS's noding, and notes that the 22 rows were not individually inspected. The raster arm is the third method: rasterise the strokes and take `skimage.measure.euler_number`. On every eighth draw (66 of the 528) it gets 2 wrong at 512 px and 7 wrong at 4096 px, so the higher resolution does worse. `euler_number` is exact and linear, and the repo makes no speed claim against it. The point is that it needs a resolution, and its answer moves with that resolution. (Note: The commits RESULTS.md names as its machine of record (`5713012`, `6d63c43`, `4730829`, `2db2d8c`, and `4a7a3dc` for real files) are not in the depth-1 clone I read at `564427f`. Every number in this essay is the repo's own reported result. I did not rebuild, re-run or re-benchmark anything, and that includes the "500 passed" test count.) Running it on your own drawing is the plainest way to see what it does. The next figure runs a browser port of the pipeline on straight line work you draw or paste. You get planimeter's full report (integers, snap window, ratio, radius, number of candidate windows tried) or its refusal block, and next to it the round-6 strawman's integer for the same input. The presets include the README's 2×2 grid. Its window there is $[1.819\times10^{-12},\,1)$, with ratio $5.498\times10^{11}$ and radius $1.349\times10^{-6}$, and the line printing it is the one the README calls "the line no other tool prints". **Figure 2.** Draw or paste straight line work (line, polyline, polygon, rect, and path with M/L/H/V/Z only; curves need the CLI). The panel shows exactly what planimeter would print: pieces, faces, chi and the window that decided them, or a refusal with reason, coordinate, owning element and action. The round-6 strawman's integer for the same input is shown beside it. With lift on, each enclosed face counted by the certified arrangement rises as a slab. The slab height means only 'counted'. A face count is arrangement-theoretic, not what a renderer fills. A Predict faces box checks a stated count against the measured one. Colour key: structure: your segments and their vertices; proof: enclosed faces counted under a certified window; baseline: round-6 strawman integer, where it differs; parameter: the endpoint you hold, and the face count you predict; refused: refused: the tool names the coordinate and returns no integer. The tool also checks claims. `planimeter check plan.svg --faces 5` returns `HELD` (exit 0), `BROKEN` (exit 1) or `REFUSED` (exit 2). Three states, not two, so that "the tool could not answer" never reads as "your prediction was wrong". `init` adds two lines to `CLAUDE.md` asking the agent to state its predicted count before an edit. The README calls that "a nudge, not a result", and whether agents follow it is unmeasured. The hook is built to stay installed, and the README lists four contract clauses: silence is the default; bounded work or a budget refusal, never a stalled turn; any exception exits 0 with empty stdout; and it never writes to a geometry file. A refusal stamps a line only when its reason changes, because "Thirty identical refusal lines in a session train an agent to ignore the tool." The suffix test runs before any package import, so a write to `notes.py` costs about one interpreter start and prints zero bytes. **Figure 3.** Left: a scripted editing session. Each write shows the line the hook stamps: nothing for a .py file, a face count, a face delta between two certificates of the same window, a refusal token, and silence when the same refusal repeats. Right: the repo's wall-clock rows (20 cold subprocesses per row, Windows) with the bare-interpreter floor, a control row with the suffix check removed, and the eleven silent-path medians against the 60 ms gate. Three of the eleven exceed it. Colour key: proof: a certified stamp; measured: wall-clock medians and ranges, as reported; baseline: the floor: a bare interpreter with no hook; refused: a refusal stamp; silent-path runs over the 60 ms gate are ringed. **Measured: 56.4 ms** hook, silent path (.py write), median \[46.7, 62.9]; geometry path (.svg) 158.6 ms \[145.0, 168.7]. Control: FLOOR: bare interpreter with no hook, 52.4 ms \[48.0, 70.8]; the same hook with the suffix check removed, 154.0 ms \[138.9, 163.4]. n = 20 cold subprocesses per row. Source: RESULTS.md:267-272 @ 564427f; Windows run of record as reported. What these rows support is a range: the silent path costs "under about 15 ms" over a bare interpreter, and the suffix check saves 84 to 105 ms per non-geometry write across seven runs. The stamp body is 7 `cl100k_base` tokens, plus 5 for the path. On files nobody in the repo drew, there is no answer key. The repo therefore reports refusal rates and agreement between methods, never accuracy. On 13,681 npm icon files (tabler, feather, bootstrap), planimeter answers 10,862 (79.4%) with a median of 37 ms. On 30 icons committed to the repo, it answers all 30. The round-6 snap gives a different integer on 7 of those 30, and `skimage.euler_number` at 4096 px, a different library working on a different representation, agrees with planimeter's integer on all 7. ## How it was made The pipeline is a fixed sequence, and no step creates a coordinate: exact dedup → Euclidean minimum spanning tree → candidate windows → cluster → subdivide at *existing* vertices only → dedup edges → check no remaining crossing → count. > **Definition: The three integers.** > > Let the certified arrangement have $V$ vertices (endpoint clusters), $E$ distinct edges after subdivision and deduplication, and $C$ connected components. Then **pieces** $= C$, **faces** is the number of bounded faces, and $\chi = V - E$. The convention is arrangement-theoretic: two overlapping filled rectangles, with their crossings written into the file, are 3 faces even though a renderer shows 2 regions. ### The counting layer For a graph drawn in the plane, where $F$ counts every face including the single unbounded one, Euler's identity is $$ V - E + F = 1 + C . $$ **Figure 4.** A small planar graph on a 4×4 lattice. Click edges to toggle them. V, E and C come from the edge set (C by union-find), F = faces + 1, and the panel substitutes the live values into V − E + F = 1 + C. Each bounded face rises as a slab when a cycle closes. A chord inside a component adds one face and lowers chi by one. A bridge between components removes one piece, leaves faces unchanged and lowers chi by one. The presets reproduce the closed-form table in count.py. Slab height means only 'counted'. Colour key: structure: vertices and edges; proof: bounded faces counted by the identity; parameter: the edge you toggle. Write $\mathrm{faces} = F - 1$ for the enclosed faces. The three printed integers then follow in two lines of algebra: $$ \mathrm{faces} = E - V + C, \qquad \chi = V - E, \qquad \mathrm{pieces} - \mathrm{faces} = C - (E - V + C) = V - E = \chi . $$ **Measured: 44 / 44** closed-form families on the clean stratum counted exactly, with n_merged = 0 on every one. Control: each recipe's faces re-derived independently as E − V + C in the test file (test_every_recipe_agrees_with_euler). n = 44 families. Source: RESULTS.md:323 @ 564427f; run of record as reported. A triangle has $V = 3$, $E = 3$, $C = 1$, so faces $= 1$ and $\chi = 0$. In code it is the end of `count()` (`planimeter/count.py:106-111` @ `564427f`): ```python v = n_vertices e = len(edges) pieces = dsu.n_sets faces = e - v + pieces return Counts(v=v, e=e, pieces=pieces, faces=faces, chi=v - e, dangles=sum(1 for d in degree if d == 1)) ``` The docstring says what this layer can and cannot claim: "Given a correct edge list this file cannot be wrong, and nothing in it is evidence that the edge list is correct." The rest of the package exists to make the edge list trustworthy. ### The merge spectrum Exact duplicates come first. Bitwise-equal coordinates are one point, with `-0.0` normalised to `+0.0`. This is an identification, not a tolerance. Without it, a clean file with every corner written exactly would have no cluster structure, and the gap search would reach into the drawing's own length distribution. Write $\pi(r)$ for the partition of the remaining endpoints into groups at tolerance $r$, meaning the connected components of the graph with an edge wherever $d(p,q) \le r$. > **Theorem: Gower & Ross (1969), and Kruskal's cut property.** > > The single-linkage merge heights of a finite point set are exactly the sorted edge weights of its Euclidean minimum spanning tree. So $\pi(r)$ can change only at those $n-1$ radii. If $w_i \lt w_{i+1}$ are consecutive sorted weights, $\pi(r)$ is constant for every $r \in [w_i, w_{i+1})$, and the minimum distance between two distinct groups of that partition is exactly $w_{i+1}$. The whole spectrum of radii worth considering is therefore $n-1$ numbers, computed once: $$ h_1 \le h_2 \le \dots \le h_{n-1}, \qquad \{h_k\} = \operatorname{sort}\bigl(\,|p_a - p_b| : (a,b) \in \mathrm{EMST}(P)\,\bigr). $$ **Measured: 1e-12** relative agreement of the vectorised Prim EMST with a scalar Prim written longhand; the certified window's partition identical to an explicit all-pairs single-linkage cut on 200 sets. Control: independent implementations: test_emst_matches_bruteforce, test_snap_matches_all_pairs_single_linkage. n = 500 point sets (general, collinear, duplicate-heavy, near-degenerate). Source: RESULTS.md:320-322 @ 564427f; run of record as reported. The tree is built by all-pairs Prim, vectorised, $O(n^2)$ and exact. The design first specified GEOS's Delaunay triangulation plus Kruskal, relying on the fact that the EMST is a subgraph of a Delaunay triangulation. Measured on 793 point sets, `shapely.delaunay_triangles` under GEOS 3.13.1 left out a true EMST edge on one of them. That set was in the clustered stratum, which is exactly the kind of input this tool exists for, and the resulting tree was 25.7% heavier. A spectrum from a wrong tree certifies a scale the drawing does not have, so Delaunay was dropped, and shapely and GEOS left the runtime dependencies with it. The spectrum is more than the merge heights, and `snap.py` records that both of the following were bugs it once had. First, it includes every vertex-to-non-incident-segment distance. For segment $AB$ and vertex $p$, $t = \operatorname{clip}\!\big((p-A)\cdot(B-A)/|B-A|^2,\,0,\,1\big)$, $c = A + t(B-A)$, $d = |p - c|$, with incident pairs masked out. The face count turns on vertex-to-edge incidence, so a gap search over vertex pairs alone cannot see a wall that misses its floor by $2.2\times10^{-6}$. It also invents a doubt that is not there: in a triangle with sides of 10, the smallest vertex-pair separation is a side, so the apex, 8.66 from its own base, would read as ambiguous. Second, the spectrum starts at a representability floor $\delta = 4096 \cdot 2^{-52} \cdot m$, where $m$ is the largest absolute coordinate. Separations below $\delta$ are float noise and merge unconditionally. ### The window A candidate window is a pair of consecutive spectrum values, and it is certified when $$ \pi(r) \text{ is constant for every } r \in [t_{\text{below}},\ t_{\text{above}}), \qquad \frac{t_{\text{above}}}{t_{\text{below}}} \geq \rho, \qquad \rho = 10 . $$ **Figure 5.** The drawing's endpoints lie on the base plane. Above them, each EMST merge is drawn at height log10 of its weight, so the single-linkage dendrogram stands on the points. The representability floor (4096 ulps of the drawing's magnitude) is the base slab, and vertex-to-edge distances are ticks on the side ruler. The chosen window is a slab spanning \[log t_below, log t_above), solid when its ratio is at least rho and hatched when below. Drag the radius handle through the window to see the partition stay constant, checked against a brute-force component count at four radii across the window (toggle). The presets include a jittered 2×2 grid, two squares a gap g apart, a triangle, and a no-scale point set. Height is log distance, not drawing height. Colour key: structure: endpoints and the dendrogram; proof: a window with ratio at least rho; parameter: the radius r, the jitter sigma/g and the gap g you move; refused: refused: below rho, or under the representability floor. By the cut property, the closest two groups are exactly $t_{\text{above}}$ apart, so the grouping cannot change until the tolerance reaches the top of the window. Because the window is at least $\rho$ wide, "the grouping survives a tenfold change in the tolerance." The README is precise about what that does not mean: "That is strictly weaker than "the grouping is right", and saying which of the two is claimed is the entire point of the tool." `snap.py` also gives the reason no separation check is run: "a predicate that cannot fail is not evidence." The candidate list is short, and its order matters (`planimeter/snap.py:307-312` @ `564427f`): ```python lo, hi = s[:-1], s[1:] ratio = hi / lo ok = np.nonzero(ratio[1:] >= rho)[0] + 1 # genuine gaps only order = list(ok[np.argsort(-ratio[ok], kind="stable")]) if ratio[0] >= rho: order.append(0) # merge nothing, last ``` Genuine gaps are tried widest ratio first. The gap between the floor and the smallest real separation, which I call the merge-nothing window, is offered last. Its ratio is huge for a purely arithmetic reason: the floor is a few thousand ulps of the drawing's magnitude. Ranked with the others it would always win, and "every jittered endpoint is its own vertex" would become the preferred reading of every file. At most `CAND_MAX = 4` windows are kept. A window that would put more than `CLUSTER_MAX = 16` endpoints into one vertex is dropped. The radius printed with an answer is the geometric mean $\sqrt{t_{\text{below}}\,t_{\text{above}}}$, the point furthest from both edges of the window on a log scale. ### The preconditions, and the refusal policy A window that passes the ratio test still has to survive the arrangement. `_try_window` checks, in this order: - **Margin.** $t_{\text{above}} \gt 64 \cdot 2^{-52} \cdot M$, where $M$ bounds the coordinates. Above this margin, every float64 orientation and incidence sign the pipeline computes is correct, so no rational arithmetic is needed. - **P1, every edge survives.** No input segment may have both endpoints in one cluster. - **P2, robustness.** Every (vertex, non-incident edge) pair must sit at distance exactly zero or at least $t_{\text{above}}$: $$ d(p, e) \in \{0\} \,\cup\, [\,t_{\text{above}},\ \infty) \quad \text{for every vertex } p \text{ and non-incident edge } e . $$ **Figure 6.** The six conditions behind CERTIFIED, run in code order on nine fixtures (two hold all six; six each fail exactly one check, none of them check 5; one is a budget refusal): a window with ratio at least rho exists; every segment survives the merge (P1); every vertex is at exactly 0 or at least t_above from every non-incident edge (P2); no two segments still cross after subdivision (P3); the float64 margin holds; and curves flattened at N and 2N give the same integers. The run stops at the first failing card and rings the site on the drawing with its element id and distance. A budget refusal appears as a separate card labelled 'machine, not drawing'. Colour key: structure: the fixture drawing; proof: a check that passed; withdrawn: the check that failed, with its reason code. - **P3, plane embedding.** After subdivision and edge deduplication, no two segments may intersect except at a shared vertex. P2 is the check the tool turns on. A distance of exactly zero means the vertex already lies on the edge, so the edge is subdivided there. No coordinate is created, because the point already exists in the file. A distance strictly between zero and $t_{\text{above}}$ is the ambiguous band, and it refuses (`planimeter/arrange.py:407-421` @ `564427f`): ```python band = np.nonzero(d > 0.0)[0] if len(band): rows = [{"xy": [float(R[vi[k], 0]), float(R[vi[k], 1])], "element": _vertex_owner(vi[k], ends, labels, ids), "edge": owner[ei[k]], "d": float(d[k])} for k in band] rows.sort(key=lambda r: r["d"]) sites, more = _sites(rows) return Refused(REASON.VERTEX_NEAR_EDGE, detail="%d vertex%s inside the ambiguous band (0, %g)" % (len(band), " sits" if len(band) == 1 else "es sit", w.t_above), look_at=sites, n_more=more, action="move this end onto the wall, or away from it by more than %g" % w.t_above, source=src) ``` P3 is a strict four-sign orientation test. P2 has already removed endpoints lying on edges and collinear overlaps, so whatever is left is a proper crossing, and by the margin lemma its signs are correct in float64. planimeter never computes an intersection point. When two segments cross at a point the file does not contain, the answer is `EDGES_CROSS` plus the crossing's location, given for information only, never inserted into the graph. What happens after a failure is a policy, and the code states it. `EDGE_COLLAPSED` and `MARGIN_TOO_SMALL` mean the *window* is wrong for this drawing, so the next window is tried. `VERTEX_NEAR_EDGE` and `EDGES_CROSS` mean the *drawing* is ambiguous or unnoded. A finer window would find a reading where the near miss counts as a clean miss, and the comment in `arrange.py` says what returning that would be: "choosing the reading that happens to certify". Those refusals end the run. Curves have no vertex set, so flattening them invents one. A file with curves is run twice, at $N = 16$ and $N = 32$ samples per curve, and the verdict stands only if (status, pieces, faces, χ) agree; otherwise it refuses `CURVE_UNSTABLE`. `read.py` calls this "Evidence, not a theorem". A budget refusal on the second pass is reported as a budget refusal, not as an unstable curve. Every refusal carries a `kind`: `geometry` means re-observe the drawing, `budget` means this machine ran out of room. The default ceiling is `BRUTE_MAX = 2000` distinct vertices, and `--max-vertices` raises it. There is no Lean in this repository. The guarantees above are a cited theorem (Gower and Ross, with Kruskal's cut property), two lemmas stated in `arrange.py` (the float64 margin, and that deduplication is exact under P3), and measurements against independent implementations, listed in RESULTS.md §7. Each of the seven working modules (`snap`, `arrange`, `count`, `read`, `result`, `hook`, `cli`) has a runnable self-check, for example `python -m planimeter.snap`. ## What's new in it The usual way to decide which endpoints are the same is to choose a tolerance and apply it: round coordinates to six decimals and hash them, or hand GEOS a grid through `shapely.set_precision`. The tolerance is the caller's guess, it is not printed with the answer, and the answer always comes back as an integer. On the stratum above, that approach gives 350 and 299 wrong integers out of 528. planimeter derives the tolerance from the drawing, prints the window $[t_{\text{below}}, t_{\text{above}})$, its ratio and whether it was `derived` or supplied by the user, and refuses when the drawing does not support one. The usual way to count faces is `shapely.polygonize_full`. It is faster and more general, and, as the README says against an earlier draft of its own claims, it has a real diagnostic: it returns polygons, cut edges, dangles and invalid rings. What it does not return is which identification of endpoints produced the count. On the jitter stratum it gets 336 wrong, and its first wrong answer is the T-junction at $\sigma/g = 10^{-7}$: truth 2, answer 0. The usual way to avoid choosing a tolerance is to rasterise and count holes. That swaps the tolerance for a resolution, and the answer moves with the resolution: 2 wrong at 512 px and 7 at 4096 px on the same 66 draws. On the clean families, planar K4 disagrees with the arrangement at both resolutions, because 8-connected ink closes a background pocket at each shallow junction (5 holes against 3 bounded faces), and `theta(4)` disagrees only at 512 px. The usual way to make geometry exact is exact predicates, as in CGAL. The README credits CGAL with doing that "properly, for decades". Exact arithmetic answers the exact question, though. It makes two nearly coincident vertices genuinely distinct and returns the exact face count of a graph nobody intended to draw. planimeter is asking the intended question, and its answer is either a scale with stated stability or a refusal. The free parameter in all of this is $\rho$, and the repo published what moving it does rather than asserting that a higher $\rho$ is safer. **Figure 7.** Left: the repo's rho sensitivity table, verbatim. Right: small multiples recomputed in the browser for rho = 3, 10 and 100 across the six jitter levels, on a subset of families. The live grid uses this figure's own random draws (its own PRNG, not numpy's stream), so its counts will not match the repo's. Below: one drawing (triangle, T-junction or chain(2)) at sigma/g = 1e-2, its spectrum strip with every candidate window, and the window the pipeline chose at the current rho. At rho = 100 the genuine merge window is gone and the merge-nothing window certifies the triangle as three disjoint segments. There is no refusal cliff in either table, and the figure does not draw one. Colour key: proof: exact integers; withdrawn: planimeter's own wrong integers; parameter: rho, the minimum window ratio; refused: refused, or a window below rho; measured: the repo's verbatim counts (wrong / exact / refused). **Measured: 3 wrong** planimeter at rho = 100: 444 exact, 81 refused. At rho = 3 and rho = 10: 0 wrong, 495 exact, 33 refused.. Control: truth by construction; the three wrong rows are T-junction, chain(2) and triangle at sigma/g = 0.01, all certified with merged 0. n = 528. Source: RESULTS.md:169-176 @ 564427f; run of record as reported. Raising $\rho$ does not make the tool stricter. It removes the genuine merge window, whose ratio the drawing limits, and leaves the merge-nothing window, whose ratio is drawing scale over machine epsilon. The repo's conclusion is "Zero-wrong is a claim about `rho <= 10`", and `bench.py` prints that paragraph under its table on every run. ## What no one else built Each ingredient here has a close relative in earlier work, so I will name them before saying what differs. - **Single-linkage clustering over a minimum spanning tree.** [Gower and Ross (1969)](https://doi.org/10.2307/2346439) showed that the MST carries all the information in a single-linkage analysis, which is the theorem the window rests on. [Zahn (1971)](https://doi.org/10.1109/T-C.1971.223083) clustered point sets by deleting "inconsistent" MST edges, ones much longer than their neighbours. [Mojena (1977)](https://doi.org/10.1093/comjnl/20.4.359) evaluated stopping rules that choose where to cut a dendrogram from its fusion levels. All three are about where to cut and which clusters result. Whether the cut is stable under a stated factor, and a check that the resulting graph is a valid plane embedding, are not what these titles address. - **Snap rounding.** [CGAL's 2D Snap Rounding](https://doc.cgal.org/latest/Snap_rounding_2/) turns an arbitrary-precision arrangement of segments into a fixed-precision one by rounding to the centres of a pixel grid. [Iterated snap rounding (Packer and Halperin, 2002)](https://www.cgl.cs.tau.ac.il/projects/iterated-snap-rounding/) additionally guarantees that every vertex is at least half a pixel from every non-incident edge. That guarantee is the closest relative of P2. The difference is in the direction of the work. Snap rounding moves geometry to the grid until the separation holds, so it creates coordinates and always returns an arrangement. planimeter never moves a point. It checks whether the drawing already satisfies the separation at a scale read from the drawing itself, and refuses with the offending coordinate when it does not. - **GEOS noding and snapping.** [GEOS's `SnappingNoder`](https://libgeos.org/doxygen/classgeos_1_1noding_1_1snap_1_1SnappingNoder.html) snaps vertices and intersection points together within a tolerance. [`OverlayNGRobust`](https://libgeos.org/doxygen/classgeos_1_1operation_1_1overlayng_1_1OverlayNGRobust.html) retries a failing overlay with a heuristic tolerance and then larger ones, up to a limit. That is the opposite of planimeter's fall-through policy. GEOS escalates the tolerance until something certifies. planimeter refuses on `VERTEX_NEAR_EDGE` precisely because trying another scale would be "choosing the reading that happens to certify". [`shapely.set_precision`](https://shapely.readthedocs.io/en/2.1.2/reference/shapely.set_precision.html) rounds to a caller-given grid, and [`polygonize_full`](https://shapely.readthedocs.io/en/2.1.2/reference/shapely.polygonize_full.html) returns the faces plus diagnostics, but not the identification behind them. - **Robust predicates.** [Shewchuk's adaptive-precision predicates](https://people.eecs.berkeley.edu/~jrs/papers/robust-predicates.abstract) get exact orientation signs from float inputs by escalating precision only when the cheap estimate is uncertain. [CGAL's 2D Arrangements](https://doc.cgal.org/latest/Arrangement_on_surface_2/) build exact arrangements on exact kernels. Both make the signs of predicates correct for the coordinates as given. planimeter needs correct signs too, but it gets them from a margin lemma: every predicate it evaluates is at least $t_{\text{above}}$ from degeneracy, and $t_{\text{above}}$ must exceed 64 ulps of the drawing's magnitude. Its real question comes before the predicates. Which coordinates did the author *mean* to be equal? - **Floor-plan vectorisation and parsing.** [Raster-to-Vector (Liu, Wu, Kohli and Furukawa, ICCV 2017)](https://openaccess.thecvf.com/content_iccv_2017/html/Liu_Raster-To-Vector_Revisiting_Floorplan_ICCV_2017_paper.html) predicts junctions with a network and assembles wall primitives by integer programming. [CubiCasa5K (Kalervo et al., 2019)](https://arxiv.org/abs/1904.01920) provides a 5,000-plan dataset and a multi-task parsing model. These start from images and learn which junctions exist. planimeter starts from a vector file that already states its coordinates, uses no training data, and returns either an exact count with its scale or a refusal. It does not return a learned reconstruction. What I can say is mine is the combination, and specifically two choices. First, the merge scale is read off a gap of ratio at least $\rho$ in a spectrum that contains vertex-to-*edge* distances as well as vertex-to-vertex merge heights, and that spectrum is what makes a near miss visible at all. Second, when no gap exists, or a vertex sits strictly inside $(0, t_{\text{above}})$ of an edge, the output is a typed refusal carrying the coordinate, the owning element and one action. It is never an integer, and it is never a coordinate the tool made up. I have not found this combination in the work above. Absence from a search is not evidence of absence, and I do not claim it is the first. The README is also explicit that the clustering rule is about ten lines of numpy that any agent could write: "The guarantee is a policy, `refuse rather than guess`, not a barrier." The README credits two of my other repositories, cleave and [sigmoid](/sigmoid), as where the widest-representable-gap rule transplants from. The two figures below check the claim as far as a browser can. **Figure 8.** The jitter stratum replayed live: the 44 families, six jitter levels and two seeds, run through a browser port of planimeter and through the round-6 strawman, each scored against the construction recipe. The left bars are the repo's table, verbatim, including the shapely, networkx and raster arms the browser does not run. The right bars are this browser's run. The replay uses its own PRNG (mulberry32 seeded from a CRC of family, level and seed), not numpy's stream, so these are a different 528 drawings and the counts will not reproduce the repo's exactly. If the port ever returns a wrong integer, the draw is listed by family, level and seed. The default seed is fixed and shown, and it is never re-rolled (it can be edited). Colour key: proof: planimeter exact; baseline: wrong integers from the strawman and other arms; withdrawn: planimeter's own wrong integers, if any; measured: the repo's reported counts; line-2: exact integers from the other arms; refused: refused: no integer returned. **Figure 9.** One of five family drawings, with optional jitter, rasterised at 128 to 2048 px with 1-px Bresenham ink. Holes are counted as 4-connected background components that do not touch the border, and that count sits next to the arrangement's certified face count. The plot shows raster faces against resolution, with the arrangement's integer as a flat line. This is a different rasteriser from the repo's skimage arm, so its disagreements need not match the repo's (planar K4 at both 512 and 4096 px, theta(4) at 512 px only). The repo's counts (2 of 44 families at 512 px, 1 of 44 at 4096 px) are printed in the readout. A pointer lens magnifies a 13 x 13 pixel patch of the raster. Colour key: proof: arrangement face count, and holes that match it; baseline: raster hole count, where it differs; parameter: resolution in pixels. ## Limitations The advertised use case is a floor plan, and on real floor plans the tool almost never answers. **Measured: 0 of 96** Wikimedia Commons floor plans answered at the default vertex ceiling (2,000); 1 of 96 at --max-vertices 6000. Control: the same tool on 13,681 npm icon files answers 10,862 (79.4%); refusal rates do not compose across corpora or ceilings. n = 96 plans fetched by a pinned recipe; no ground truth. Source: RESULTS.md:575-583 @ 564427f; realdata.py at 4a7a3dc as reported. At the default ceiling, 88 of the 96 plans refuse on budget (`TOO_MANY_VERTICES`), 4 refuse `VERTEX_NEAR_EDGE` and 4 refuse `EDGES_CROSS`. Raising the ceiling to 6,000 admits 31 more files: 30 of them arrive at a geometry refusal and 1 at an answer. "Raising the ceiling does not buy answers on this set, it buys coordinates." On four files checked at `--max-vertices 25000`, where the budget cannot bind, it is 0 of 4. The median plan has 6,398 distinct endpoints. The one plan that certifies, `Akori_church_plan.svg`, gives pieces 13, faces 13, χ 0 at radius $7.03\times10^{-8}$ with a window ratio of 20,218. The repo's own diagnosis is "A published floor plan is drawn, not built": walls end near walls, and walls cross where no corner is recorded. Those are exactly the two conditions the tool refuses. **Figure 10.** Five tabs. Real files: answered and refused counts for icons and Commons plans at each ceiling, with budget refusals kept visibly separate from geometry refusals, and the plans' endpoint-count percentiles against the 2,000 ceiling. Cost: total ms against segment count for grid(k), log-log, with the committed gates (exponent 1.3; 50 ms at 100,000). Arrangement versus renderer: two overlapping rectangles are 3 arrangement faces and 2 visible regions when noded, and EDGES_CROSS when not. The contested row: two clean squares 1.00 apart read as one piece and 1.01 apart as two. Withdrawn: each claim with the measurement that killed it. All tables are verbatim from RESULTS.md. Only the cost-curve fit, the rectangles and the contested row are computed live. Colour key: proof: answered; refused: geometry refusal (hatched): the drawing is at fault; line: budget refusal (pale fill): the machine is at fault; measured: repo-reported wall clock and counts; structure: the fitted line, rectangles and rails; withdrawn: a withdrawn or NOT EARNED claim. The repo withdraws its own claims in the same voice it uses for the results: **What failed.** - Withdrawn: Usable as a hook on a real floor plan. Killed by: Fitted cost exponent 1.83 against a committed gate of 1.3; extrapolated 740,531 ms at n = 1e5 against a 50 ms gate (RESULTS.md §5). 0 of 96 Commons plans answered at the default ceiling.. - Withdrawn: A clean refusal cliff: everything certifies at small jitter, everything refuses two levels later. Killed by: Refusals sit between 3 and 10 of 88 at every jitter level (6, 4, 4, 10, 3, 6) and are not monotone; wrong stays 0 at every level.. - Withdrawn: Raising rho makes the tool stricter. Killed by: At rho = 100: 3 wrong, 444 exact, 81 refused. Zero-wrong is a claim about rho <= 10.. - Withdrawn: GEOS Delaunay plus Kruskal gives the exact EMST. Killed by: 1 of 793 point sets (clustered stratum) omits a true EMST edge; the tree is 25.7% heavier.. - Withdrawn: A CURVE_UNSTABLE refusal on 11 of 96 plans meant the curve was ambiguous. Killed by: The 2N pass had hit the vertex ceiling. Relabelled as a budget refusal; one of those plans, at 1,295 vertices, certifies.. - Withdrawn: Moving the ceiling check before the vertex-edge pass speeds up the default run. Killed by: Guard-first 361 ms median, spectrum-first 449 ms, guard-first repeated 443 ms: the drift exceeds the difference.. - Withdrawn: The vision comparison (G4). Killed by: Cut: it read coordinate text, not a render.. **Cost.** The comfortable range is under roughly 600 segments: 43.3 ms at 544 segments, 97.2 ms at 840, and 1,880.1 ms at 3,784. The cost is the all-pairs (vertex, edge) pass. A k-d tree does not remove it, because the certificate quantifies over every such pair. On one real 22,261-vertex plan, the spanning tree took 4.50 s and the vertex-edge spectrum took 40.01 s. A sweep line is on the roadmap, with the all-pairs pass kept as its test control. **Baselines that have not run.** The strongest comparison arm is a hand-written round-6 script that the benchmark itself labels `STRAWMAN`. RESULTS.md says it "is not the baseline the headline may be quoted against". That baseline is G1: twenty unprompted, first-attempt agent scripts with the whole distribution published. It has not been collected, and if its median clears roughly 20 of 24 on the stratum, "the accuracy headline is dead". The agent-written corpus with independently established truth (G5 and G6) has not been collected either. `corpus/found/` is empty. No agent-written file has been measured. **The contested row.** Two clean 10-unit squares 1.00 apart certify as one piece. At 1.01 apart they certify as two. There is no jitter in either file. The boundary is gap $=$ feature$/\rho$. This is inside the stated convention and outside what anyone looking at the picture would say. The repo also does not establish that the certified scale is unique. At most four windows are tried, the first one that passes wins, and a later one might have passed with a different answer. **What the count is not.** `faces` is arrangement-theoretic and will not match what a renderer fills. Refinement stability across $N$ and $2N$ is evidence, not a theorem: two nearly tangent curves can gain or lose an intersection at any positive flattening tolerance. The tool reads SVG line work only. The hook relies on someone else's `PostToolUse` payload schema, and a change there silences it. The 60 ms silent-path gate sits 8 ms above the bare-interpreter floor and failed on 3 of 11 runs, which RESULTS.md calls "a coin flip on process startup". All results come from one machine, one OS and one Python (Windows 11, 3.11.9, numpy 2.4.6). **Chosen numbers.** The README and RESULTS.md describe three constants the package chose: $\rho$, `CAND_MAX` and `FLOOR_ULPS`. `snap.py` also labels `CLUSTER_MAX = 16` as policy, which makes four, next to the budgets `BRUTE_MAX = 2000` and `VE_BUDGET`. The README also says the derived window involves "no tolerance you invented", while listing $\rho$ as a choice. All of these are scale-free, printed on every answer and movable by flag, and choosing them is still choosing. **Inconsistencies between the repo's documents.** A comment in `corpus.py` says that at $\sigma/g = 5\times10^{-2}$ "six draws of 108" come back with a different integer. RESULTS.md measures 4 wrong of 132 at that level (44 families × 3 seeds). A comment in `read.py` gives the curve-relabel count as 5 of 96 plans, where RESULTS.md gives 11. RESULTS.md §2 says refusals sit at 3 to 10 of 88 per jitter level, and §11 restates it as 5 to 10. I report these as they stand and quote the tables' numbers. **Provenance.** The commits RESULTS.md cites as its machine of record are not in the depth-1 clone this essay was written from, so every number here is the repo's reported result, and none was reproduced. The "500 passed" test count was not re-run. The live figures in this essay run their own port of the package with their own random draws, so they will not reproduce the repo's exact counts, and their captions say so. ## Read more [View the project](https://github.com/teerthsharma/planimeter) · [Source on GitHub](https://github.com/teerthsharma/planimeter) - Repository: [github.com/teerthsharma/planimeter](https://github.com/teerthsharma/planimeter). There is no project site; the button above goes to the repo. - Short link on teerth.dev: [teerth.dev/planimeter](https://teerth.dev/planimeter). - [README.md](https://github.com/teerthsharma/planimeter/blob/564427fb8843c24894468cec9d939af4d15ca73e/README.md): the exact statement, the refusal codes, the hook contract, prior art and "What we got wrong". - [RESULTS.md](https://github.com/teerthsharma/planimeter/blob/564427fb8843c24894468cec9d939af4d15ca73e/RESULTS.md): every table with its control, the arms that lost, NOT EARNED and NOT RUN, and the real-file runs. - [planimeter/snap.py](https://github.com/teerthsharma/planimeter/blob/564427fb8843c24894468cec9d939af4d15ca73e/planimeter/snap.py): the window, the theorem it rests on, and the policy constants. Its docstring calls it "THE FILE A REVIEWER OPENS FIRST". - [planimeter/arrange.py](https://github.com/teerthsharma/planimeter/blob/564427fb8843c24894468cec9d939af4d15ca73e/planimeter/arrange.py): P1–P3, the margin lemma and the fall-through policy. - Related essays: [tangle](/tangle) and [separatrix](/separatrix), the other tools here that certify an integer or refuse; [topological-ml-toolkit](/topological-ml-toolkit), for single-linkage clustering and H0 in another setting; [sigmoid](/sigmoid), named in the README as one source of the widest-gap rule. --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # Epsilon-Hollow > A research kernel that keeps its state as points on a sphere and returns each discrete answer with a certificate that rounding could not change it, or refuses. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/epsilon-hollow - Repository: https://github.com/teerthsharma/Epsilon-Hollow - Project site: https://teerth.dev/Epsilon-Hollow/ - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "Epsilon-Hollow", 2026, https://teerth.blog/epsilon-hollow - Languages: Rust, Lean 4 - Question: Can a kernel read the shape of the ML work it runs, and refuse any integer that floating-point rounding, not the data, decided? - Headline result: 1,215 → 0 wrong nearest-centroid answers in 5,000 seeded queries, before and after the certificate (control: reported hit rate 1.0000 before, 0.9818 after; stated on the project page, not re-measured) ## What it is Epsilon-Hollow is the repository; the kernel inside it is called Seal OS. It is a research kernel for x86_64, written in `no_std` Rust and booted only as a UEFI application. A conventional kernel measures its workload by quantity: resident pages, CPU time, open files, queue depth. It keeps no model of the workload's shape. To Linux, a training run is a process with a large heap and an inference server is a process with a larger one. Whether the run is converging or memorising its training set, and which prefixes live conversations share, are facts the kernel could compute and never does. Seal OS starts from the other end. Physical frames, files and tasks are placed on the unit sphere $S^2$ as points or small point clouds, each owned by the Voronoi cell of its nearest centroid, and placement and prefetch are read off that geometry. A trainer hands the kernel two numbers per step, and the kernel names the regime the run is in from the shape of its validation curve. An inference server's KV cache is a prefix tree in kernel memory. Shape comes back as integers: how many components, which $k$ keys, which cell, whether a curve folds back on itself. Those integers are computed in floating point, and near a decision boundary "the rounding of the arithmetic, not the data, picks the answer." So the kernel holds itself to one rule. It computes an enclosure of the quantity. If the enclosure clears the boundary, the answer leaves with a certificate. If it touches the boundary, the answer is a refusal that names the witness, and the caller widens, scans or gives up. > **Definition: Certify or refuse.** > > A discrete answer $a$ computed from data $x$ in floating point is **certified** when an enclosure $[\,\underline{f}(x), \overline{f}(x)\,]$ of the underlying real quantity lies strictly on one side of every decision boundary that separates $a$ from another answer. Otherwise the result is a **refusal** that names the input closest to the boundary. A refusal says the arithmetic cannot decide; it does not say the answer is wrong. The rule binds the kernel's own claims too. At boot the kernel evaluates ten theorem lines, T1 to T10, against the values its running subsystems were built from. The code at commit `36b1400` prints one verdict per line, `CERTIFIED`, `NOT CERTIFIED` or `NOT CHECKED`, and a tally. CI greps for this tally, and `seal-mkimage --check-theorem-log` recomputes it from the ten lines: ```text [BOOT] Theorems: 2 certified (T1/TSS T2/SCM), 1 not certified (T4/AGCR), 7 not checked (T3/GMC T5/HCS T6/RGCS T7/PHKP T8/TEB T9/CMA T10/WPHB) ``` The README and `docs/THEOREMS.md` still quote an older banner, "9 of 10 theorems VERIFIED"; the code is stricter than the docs, and I follow the code. T4 is the governor's convergence claim, and the kernel refuses it at the step the governor actually runs. Seven of the ten lines are not checked at all, because the kernel has no running instance of what each one describes. **Figure 1.** The boot theorem gate, run on a log you can edit. The default log is the fixture of a live boot from the repo's gate tests: T1 and T2 certified, T4 not certified at alpha + beta/dt = 5.01 at dt = 0.01, seven lines not checked. The gate recomputes each verdict from the evidence printed on its line and rejects a bare VERIFIED banner (the pre-live format) or a T4 line that claims certification while the margin is at least 1. Below the log, a ledger sets each theorem's boot verdict beside what the docs say and how strong its Lean artifact is. A strip shows the 96,110 lines of Rust under kernel/seal-os/src at commit 9ebbe2e by subsystem; 4,073 lines are not itemised by the repo's table and are drawn hatched. Colour key: proof: CERTIFIED verdict; gate PASS; ledger tick (certified, or a full Lean artifact); withdrawn: NOT CERTIFIED verdict; gate REJECT; ledger square (not certified, or refused in the docs); baseline: NOT CHECKED: no running instance of it in the kernel; ledger open circle (not checked, or a Lean placeholder); ink-2: ledger half circle: partly, or a layered Lean artifact; structure: lines of Rust per subsystem; shade groups subsystems; refused: lines not itemised by the repo's table. What the kernel is not matters as much. No instruction has executed in user mode; the applications are kernel code. Bring-up of the other processors is written, but its only call is commented out. Hardware coverage is one QEMU configuration. So I do not describe Seal OS as a general operating system. It boots under QEMU, brings up its drivers, mounts its filesystem, replays its ML proofs on synthetic traces and evaluates its own theorem lines, and those are the parts this essay is about. One tool grew out of this work: [separatrix](/separatrix), which has its own essay. ## What it can do The README names four pillars, and says of them: "Each runs in the kernel today, on synthetic workloads. Validation on real models is the open work." CI numbers below are from run 36165748105 on commit `9ebbe2e`, the run in which CI is red (see Limitations). ### `stratum`: a training run has a shape A trainer passes `(train_loss, val_loss)` once per step through system call 121 (`SYS_FIT_OBSERVE`); call 122 (`SYS_FIT_REGIME`) returns `Underfit`, `WellFit`, `Overfit` or `Collapsing`. The kernel sees nothing else of the model: no weights, activations or gradients. Non-finite input is counted and latched, never used. The test is geometric. The last 64 validation losses become delay points $p_t = (v_t, v_{t-1}, v_{t-2})$. A run that only falls draws an open arc. A run that falls and then climbs back through values it already visited draws a V, and at a large enough scale the two arms of the V close into loops. In the README's words: "Overfitting is revisitation, and revisitation is a cycle." The construction is in How it was made. **Measured: 7 / 7** synthetic training runs classified correctly: underfit, wellfit, overfit, collapsing, a negative control, a monotone line, a monotone exponential. Control: a train/validation gap threshold flags the negative control (a healthy run with a constant validation offset of 0.35) as overfit; stratum does not. n = 7 streams of 128 steps, window 64, κ = 1.68. Source: CI run 36165748105, commit 9ebbe2e. **Measured: 4,792 B** memory per stratum stream. Control: bounded over a 4,096-step stream. Source: CI run 36165748105, commit 9ebbe2e. The verdict is advisory. The `FitAction` documentation says so directly: "Every field is advisory today, and nothing here is enforced against a task that ignores it." `SYS_FIT_REGIME` publishes a prefetch threshold, and the only reader of that threshold is a prefetch preset that, in the source's words, nothing constructs. The figure below runs the same classifier on a curve you paste. **Figure 2.** Paste a validation-loss curve (and optionally a training-loss curve) and the stratum classifier, ported from trajectory_shape.rs and stratum.rs, names its regime, next to the train/validation gap rule on the same input. The shaded band is the 64-point window the kernel keeps. The cascade strip shows the fixed order of gates (Collapsing, then Overfit, then Underfit, else WellFit); the first lit gate is the verdict. Presets are the seven synthetic proof fixtures from stratum.rs. The thresholds were set on those seven synthetic streams; the classifier has never seen a real model, and a curve you paste is judged by the same thresholds. Colour key: structure: the 64-point window in use; proof: verdict matches the fixture's ground truth; withdrawn: verdict differs from the fixture's ground truth; first non-finite value; baseline: gap-threshold verdict on the same input; parameter: step cursor and calibration sliders. ### `foliation`: an inference cache is a prefix tree A sequence's block table is its path down a prefix tree the kernel holds. A block is 8 tokens; each resident block is one 4 KiB physical frame from the kernel allocator. Two sequences that agree on a block-aligned prefix land on the same blocks, and there is no call to share a block: appending identical tokens does it. The source says it plainly: "Prefix sharing is not a hash table bolted onto an allocator." A 64-bit digest of the prefix narrows the search, and a child is shared only when the token array also matches, because the digest can collide. Eviction may remove only a free face: a block that is resident, unreferenced, and has no resident children. The resident set therefore stays a connected rooted subtree under every policy, and policies differ only in which free face they pick. The foliation policy picks the leaf fewest sequences ever entered, then the deepest, then the oldest. The repo measures it on two synthetic traces at a 24-block pool, with "bp" meaning basis points of descents that hit a shared block: **Measured: 952 bp** foliation hit rate on the boot trace (9.52 %), equal to the Belady oracle. Control: LRU 0 bp; locality-only null 476 bp; random 619 bp at seed 0, 238 to 857 bp over 32 seeds (foliation wins on 32 of 32). n = 30 requests, 1,680 tokens, 24-block pool, 8 tokens per block. Source: CI run 36165748105, commit 9ebbe2e; locality null from commits 264235c and 0ab2377, local QEMU. That trace is built so that recency always evicts the shared prefix: the hot prefix returns only after 31 other blocks, more than the pool holds, and a `const` assertion in the source states the construction. LRU's 0 is a property of the trace. On a multi-turn chat trace, where reuse follows recency, the result flips: **Measured: 5,284 bp** foliation hit rate on the chat trace. Control: LRU 8,068 bp; Belady 8,143 bp; locality-only null 6,818 bp; random 7,026 to 7,443 bp; foliation beats random on 0 of 32 seeds. n = 16 conversations, 4 live, 6 turns each; 96 requests, 528 descents. Source: commit 0ab2377, local QEMU; random range from a host replay of the same module. With 16 live conversations instead of 4, one mutation build reverses it again (5,113 bp against LRU 3,731 and Belady 5,378). The repo's reading is the one I hold: "Which policy wins follows whether reuse distance exceeds the pool, not the name of the request shape." Because of that, the cache behind the system calls defaults to LRU, and the source comment at its constructor says the foliation ranking "beats LRU only at a capacity cliff on the synthetic boot trace and ties or loses elsewhere." **Measured: 0 / 0** referenced evictions and collapse violations on the boot trace; 20 shared descents saved 81,920 bytes. Control: 190 frames backed and 190 freed, 0 failed; every replay of either trace must also show 0 and 0. Source: CI run 36165748105; replays at commit 0ab2377. ### Certify-or-refuse, before and after **Measured: 1,215 → 0** wrong nearest-centroid answers in 5,000 seeded queries on S² (8 centroids). Control: before: reported hit rate 1.0000 while wrong; after: 4,909 certified in the 3×3 block, 91 full scans, reported hit rate 0.9818. n = a test in aether-core asserts 0 wrong over 25,000 queries at K = 2, 8, 20, 64 and 200. Source: project page, commits cdb4a4e and 0f040e0: stated, not re-measured. The matching rows for component counts and attention top-k sit next to their equations in How it was made. ### TopoRAM and ManifoldFS: state on a sphere Every physical frame carries an embedding of 16 points on $S^2$ (32 quantized angles, 64 bytes), a 64-tick access history, a Voronoi cell and a lifetime class. Memory is split into three zones, each with eight seeds; a frame belongs to the cell of its nearest seed. **Measured: 64 / 64** TopoRAM allocations that land in their target cell. Control: no fallback cell taken; p50 6,126 and p95 12,446 cycles per allocation. n = 64 allocations. Source: boot benchmark under QEMU TCG on a GitHub-hosted runner; emulated TSC, not hardware. ManifoldFS encodes a file's bytes as a point cloud on $S^2$ and files the inode in the Voronoi cell of the cloud's first point; the bytes persist through ext2. **Measured: ≤ 7 ops, 0 B** ManifoldFS same-inode move: metadata operations and bytes of file data written. Control: mock block store, fs_mode=mock_block, persistence_bytes_per_move=0. Source: checked at every boot by --check-benchmark-log; CI run 36165748105. The README adds the sentence that has to go with both numbers: "Whether this layout beats a conventional allocator or filesystem has not been measured; no comparison against Linux exists." ### Upstream Work on Epsilon-Hollow led to these fixes in two other projects. I state the link in those words and no more. **Upstream: **[openxla/xla #46539](https://github.com/openxla/xla/pull/46539), fix(gpu): make reduction group order deterministic (landed as 3d5df1d). Switches the set that drives the union-find merges in GroupDisjointReductions to absl::linked_hash_set, so the group-to-block assignment no longer depends on hash-set iteration order (as the PR states it). Work on Epsilon-Hollow led to this fix. **Upstream: **[tensorflow/tensorflow #124410](https://github.com/tensorflow/tensorflow/pull/124410), Fix transitive reduction of collective control edges (merged 2026-08-05). Closes reachability exactly in one pass, so the unique transitive reduction of the collective control edges is emitted; the PR's four-op example drops a redundant edge (four edges to three). Work on Epsilon-Hollow led to this fix. ## How it was made The kernel is one `no_std` crate, `kernel/seal-os`, built for `x86_64-unknown-uefi`. Drivers, filesystems, the network stack, the window manager and the applications compile into one EFI image. The mathematics lives in a workspace crate, `aether-core`: certified $\beta_0$, certified top-k, trajectory shape, the spherical Voronoi indices, the governor. The kernel calls into it. I follow the repo's own numbered equations, (1) to (9), then the Lean files, then the paging that holds it all. ### Certified β₀ Let $h_1 \le \dots \le h_{n-1}$ be the edge weights of the Euclidean minimum spanning tree of a point set $X$; these are exactly the single-linkage merge heights. The number of components at threshold $t$, and the condition under which a count at scale $s$ with band ratio $r \ge 1$ is certified, are: $$ \beta_0(X, t) = n - \#\{\,k : h_k < t\,\} \tag{1} $$ **Measured: 500 / 500** seeded point clouds on which the certified count agrees with all-pairs union-find. Control: before the band rule, two points at 0.5 ± 1e-9 gave two different integers; both are refused now. Source: project page, commit 8678c97; cargo test -p aether-core --test certified_betti. $$ \text{certified at } s \iff h_k \notin \left[\, s/\sqrt{r},\; s\sqrt{r} \,\right] \quad \text{for all } k. \tag{2} $$ **Figure 3.** Certified β₀ on a point cloud you edit. Panel A draws the points and their minimum spanning tree; edges below the band are merged at the scale, edges inside the band are the ones whose comparison with the scale rounding could decide, edges above it stay split. Panel B places the merge heights on a log axis with the band \[s/√r, s√r]; when no height lies inside it the band is certified. Panel C is the staircase of equation (1), with an independent union-find count laid over it. Panel D maps the (scale, band ratio) plane: certified cells carry their integer, refused cells are hatched. A two-point preset at distance 0.5 ± 1e-9 shows the naive count flipping while the certified rule refuses all three distances. Colour key: structure: points of the cloud; ink-2: below-band MST edges (merged at the scale); merge-height ticks; baseline: in-band edges: decided by rounding under a bare comparison; naive count; proof: certified band and its integer; refused: refused cells of the (s, r) plane; withdrawn: union-find disagreement (expected never); parameter: scale s and band ratio r. Every threshold in the band, under `<` or `<=`, gives the same integer, which is what the certificate means. If a height lies in the band, the result is `Refused { i, j, height }`, naming the in-band tree edge whose height is nearest $s$. Heights are computed as $m\sqrt{\sum_d (\delta_d/m)^2}$ with $m = \max_d |\delta_d|$, so separations near $10^{-170}$ or $10^{170}$ neither underflow nor overflow. `certified_beta0` builds the tree with all-pairs Prim, $O(n^2)$. The rule is ported from planimeter's gap rule. ### Certified attention top-k Each score $\hat{s}_j = \mathrm{fl}(q \cdot k_j)$ of head dimension $n$ carries Higham's a-priori bound for a floating-point inner product ([Accuracy and Stability of Numerical Algorithms](https://epubs.siam.org/doi/book/10.1137/1.9780898718027), Chapter 3; the repo cites it as Theorem 3.1): $$ \left|\hat{s}_j - q \cdot k_j\right| \le \gamma_n \sum_{d} \left|q_d\, k_{j,d}\right|, \qquad \gamma_n = \frac{n u}{1 - n u}, \quad u = 2^{-53}. \tag{3} $$ **Measured: 0.0 vs 1** float score of key 0 against its exact score, for k₀ = \[10¹⁷, 1, −10¹⁷] and q = \[1, 1, 1]; key 1 (score 0.5) was taken over key 0. Control: after the rule: the row is widened instead of answered. Source: project page, commits afd0969 and eb4af16; cargo test -p aether-core --test attention_contracts. The absolute value sits inside the sum. The source comment is specific about why: "`gamma_n · |q·k|` is not this bound and is far below it under cancellation, which is exactly where it matters." The computed radius inflates (3) by $(1 + 2\gamma_n + 4u)$, adds $n$ times the smallest subnormal, and rounds up, so it bounds rather than estimates: ```rust // kernel/epsilon/epsilon/crates/aether-core/src/attention.rs:885-898 @ 36b1400 fn enclosed_dot(q: &[f64], k: &[f64], i: usize, j: usize, head_dim: usize) -> (usize, f64, f64) { let (mut sum, mut magnitude) = (0.0f64, 0.0f64); for d in 0..head_dim { let product = q[i * head_dim + d] * k[j * head_dim + d]; sum += product; magnitude += product.abs(); } let n = head_dim as f64; let u = f64::EPSILON / 2.0; let gamma = nextafter(n * u / (1.0 - n * u), f64::INFINITY); let underflow = n * f64::from_bits(1); let radius = gamma * magnitude * (1.0 + 2.0 * gamma + 4.0 * u) + underflow; (j, sum, nextafter(radius, f64::INFINITY)) } ``` The absolute sum rides the same pass as the score, so the enclosure adds no dot products. With radii $r_j$ in hand, the top set $T$ of size $k$ is certified when every kept lower end clears every excluded upper end: $$ \min_{i \in T}\left(\hat{s}_i - r_i\right) \;>\; \max_{j \notin T}\left(\hat{s}_j + r_j\right). \tag{4} $$ **Figure 4.** Certified top-k in three panels. A: type a query and up to six keys (or use the cancellation preset q = \[1,1,1], k₀ = \[M, 1, −M] with M = 10^p); each key's float score, exact score (computed in exact rational arithmetic) and Higham interval are drawn on one axis, and the certificate either holds, or refuses and widens the row, next to the plain float argmax. B: random trials compare the actual rounding error of each dot product with its computed radius on log axes; points above the diagonal would be violations of equation (3). A halo marks trials whose error exceeds γₙ·|q·k|, the bound the source warns against. C: the rule of equation (4) on the repo's test fixture (solid boxes are the keys inside the budget, dashed boxes the rest), where comparing only rank k with rank k+1 would certify a set the full rule refuses. Colour key: baseline: float point estimate; float argmax; proof: exact score; CERTIFIED; the diagonal err = r; withdrawn: REFUSED or WIDENED; a bound violation; measured: witness vector inside the boxes; line-2: Higham interval around the float score; parameter: the bit of fl(M) with weight 2^0, which moves with the exponent p. The comparison is strict, and it compares the lowest lower end inside against the highest upper end outside, not rank $k$ against rank $k+1$: ```rust // kernel/epsilon/epsilon/crates/aether-core/src/attention.rs:652-673 @ 36b1400 let (inside, outside) = ranked.split_at(budget); let lowest = inside .iter() .copied() .min_by(|a, b| lower_end(a).total_cmp(&lower_end(b))) .expect("budget > 0"); let highest = outside .iter() .copied() .max_by(|a, b| upper_end(a).total_cmp(&upper_end(b))) .expect("budget < len"); if lower_end(&lowest) > upper_end(&highest) { Ok(inside.iter().map(|&(j, _, _)| j).collect()) } else { Err(TopKRefusal::Boundary(BoundaryRefusal { inside: lowest.0, outside: highest.0, gap: lowest.1 - highest.1, needed_margin: lowest.2 + highest.2, })) } ``` On refusal, `certified_or_widened` keeps every key whose upper end reaches the lowest selected lower end, at no extra dot products; a NaN or infinite score refuses first and the row falls back to dense. The separation test is ported from separatrix's error intervals. ### The fold score of a loss curve From the validation losses the kernel keeps the last 64 delay points $p_t = (v_t, v_{t-1}, v_{t-2}) \in \mathbb{R}^3$ and resamples them to uniform arc length. Let $\varepsilon^\ast$ be the largest edge of the resampled cloud's minimum spanning tree and $m$ the number of resampled points. The loop score is $$ \ell = \begin{cases} 0 & \text{if the window is monotone,} \\ \min\!\left(1,\; c(\kappa\,\varepsilon^\ast)/m\right) & \text{otherwise,} \end{cases} \qquad c(\varepsilon) = E_\varepsilon - V + \beta_0, \tag{5} $$ **Measured: 0.969 → 0** loop score of a monotone staircase. Control: before: scored 0.969 and judged Overfit; after: certified 0 before any complex is built. Source: cargo test -p aether-core --test trajectory_shape (monotone_staircase_scores_no_fold). where $E_\varepsilon$ counts Vietoris–Rips 1-skeleton edges at scale $\varepsilon$ except two-step chords already filled by a triangle, so $c$ upper-bounds Rips $\beta_1$ rather than equalling it. Monotonicity is checked by comparing stored values coordinate by coordinate, with no rounding, so the zero case is exact. The one free constant, $\kappa$, is bounded by a derivation rather than tuned: $$ \sqrt{8/3} \approx 1.633 \;<\; \kappa = 1.68 \;<\; \sqrt{3} \approx 1.732. \tag{6} $$ **Figure 5.** The delay cloud of a loss curve in three dimensions, with the Rips edges at scale κ·ε\* drawn between its points. Rotate it: for a clean symmetric V the falling and rising arms are two parallel rails, and the edges that close loops are the rungs between them. Below, the loop score of equation (5) is plotted against κ for the V, the staircase with and without the chord quotient, the selected fixture and two values quoted from the repo comment, with the band of equation (6) shaded; two toggles switch off the monotonicity certificate and the two-step-chord quotient, reconstructing (not replaying) the two repairs the repo made. The last panel maps the participation ratio of equation (7) over the autocovariance ratios, with the seven stratum fixtures marked and the 0.45 Underfit threshold drawn. Colour key: seq: time order of the points (colour bar); participation ratio in the equation (7) map; parameter: the κ cursor; proof: the derived κ band; baseline: readings the repairs removed; measured: the Rips scale sphere; values recomputed from the repo's code; refused: PR withheld (NaN): no point drawn. The floor comes from a symmetric V. Its descending point $(k, k+1, k+2)s$ and ascending point $(k+2, k+1, k)s$ differ by $(2, 0, -2)s$ for every $k$, so the arms first meet at $\sqrt{8}\,s = \sqrt{8/3}\,\varepsilon^\ast$. The ceiling comes from a monotone stretch: every segment of the delay polyline lies in one closed orthant, so a chord spanning $k$ resampled steps is at least $k\varepsilon^\ast/\sqrt{3}$ long, and since two-step chords are quotiented out, the first chord that can close a cycle spans three steps and needs $\sqrt{3}$. The value 1.68 is the midpoint 1.6825, rounded. The source records the measurement against it: the clean symmetric V scores 0 up to $\kappa = 1.62$ and 0.391 from 1.64. It also records, in the same comment, that the band "is proved for the symmetric V only." I come back to that under Limitations. A loop alone does not mean overfitting; a converged run sitting in a noise ball also loops. So the classifier needs a second signal, the drift of the residual $v_t - \text{train}_t$, and it reads underfitting from a third, the participation ratio of the training-loss autocovariances $c_0, c_1, c_2$: $$ \mathrm{PR} = \frac{3}{3 + 4(c_1/c_0)^2 + 2(c_2/c_0)^2} \in \left[\tfrac{1}{3}, 1\right]. \tag{7} $$ **Measured: 0.353 / 0.814** participation ratio of the underfit fixture and of the converged fixture. Control: Underfit threshold 0.45, about 35 % above the rank-1 floor of 1/3; the overfit fixture's training loss scores 0.424, which is why Overfit is tested before Underfit. Source: trajectory_shape.rs:514-517 and 683-686 @ 36b1400 (measurements stated in the source). PR is returned as NaN, a refusal, when the training loss varies by less than its own rounding bound, and the Underfit verdict is then withheld. `classify` decides in a fixed order: non-finite input or an unmeasurable signal gives Collapsing; fewer than 16 samples gives WellFit; training-loss drift of at least 0.50, or a shatter ratio of at least 100 with any rise in training loss, gives Collapsing; $\ell \ge 0.125$ with residual drift of at least 0.05 gives Overfit; a finite PR at or below 0.45 gives Underfit; anything else is WellFit. ### The governor, and why T4 is refused The governor adapts a wake-up threshold $\varepsilon_t$ from an observed deviation $\Delta_t$ with a proportional-derivative step. In code the target rate is $R^\ast = 1000$, $\alpha = 0.01$, $\beta = 0.05$, and $\varepsilon$ is clamped to $[0.001, 10]$: $$ e_t = R^\ast - \frac{\Delta_t}{\varepsilon_t}, \qquad \varepsilon_{t+1} = \operatorname{clamp}\!\left(\varepsilon_t - \alpha\, e_t - \beta\,\frac{e_t - e_{t-1}}{\Delta t}\right). \tag{8} $$ **Measured: 0.001 ↔ 10** the loop's behaviour over 2,000 ticks from ε = 0.1: a 2-cycle between the two clamp bounds. Control: the same 2-cycle at Δt = 0.01, 0.0506 and 1. Source: commit dcc35b6 (stated in RESULTS.md). T4 certifies geometric convergence when the gain margin is below one: $$ \alpha + \frac{\beta}{\Delta t} < 1 \;\Longrightarrow\; \rho = 1 - \frac{\alpha}{1 + \beta/\Delta t} \in (0, 1). \tag{9} $$ **Measured: 5.01** T4 gain margin α + β/Δt at the shipped gains and the runtime step Δt = 0.01. Control: the bound is 1; at Δt = 1, a step no runtime caller uses, the margin is 0.06. Source: kernel/seal-os/src/lib.rs:162-164 constants; boot gate since commit 3c14df0. **Figure 6.** The T4 condition and the loop it is meant to describe. Tab 1 plots the margin α + β/Δt against Δt on a log axis with the line at 1, marking Δt = 1 (0.06), the runtime step Δt = 0.01 (5.01) and Δt = 0.0506 (0.998), and rebuilds the boot line exactly as the kernel prints it. Tab 2 runs the update of equation (8) for 2,000 ticks at three step sizes: each one 2-cycles between the clamp bounds 0.001 and 10. Tab 3 takes one step of the Rust governor_step from the repo's test fixture and shows |e| rising while the gain-margin predicate holds, beside the Lean statements, which are not about that step. Colour key: proof: region where the margin is below 1; a step in which |e| descends; withdrawn: refusal; the 2-cycle trajectory; a step in which |e| rises; baseline: the retired VERIFIED-at-Δt=1 state; measured: trajectory points recomputed from the repo's update; parameter: the reader's Δt on the margin curve. The boot check is a direct evaluation of (9) at the constants every runtime caller passes: ```rust // kernel/seal-os/src/theorems.rs:208-223 @ 36b1400 /// T4/AGCR: `alpha + beta/dt < 1` at the gains and step every runtime /// governor uses. fn t4(state: &LiveState) -> Verdict { let (alpha, beta, dt) = state.gains; let margin = alpha + beta / dt; let rho = aether_agcr::contraction_rate(alpha, beta, dt); let certified = rho > 0.0 && rho < 1.0 && aether_agcr::half_life(rho).is_finite() && aether_agcr::gain_margin_stable(alpha, beta, dt); let relation = if certified { "<" } else { ">=" }; verdict( certified, format!("alpha+beta/dt={:.2} {} 1 at dt={}", margin, relation, dt), ) } ``` Until commit `3c14df0` the same theorem was certified at $\Delta t = 1$. Moving the step to 0.01 is not the whole story, though. Condition (9) treats the map from $\varepsilon$ to $e$ as unit gain; linearising (8) gives $\partial e/\partial\varepsilon = \Delta/\varepsilon^2$, about $10^6$ at the equilibrium $\varepsilon^\ast = 0.001$, and $\alpha K \approx 10^4$ alone exceeds one. No step size satisfies the condition. The repo says what follows: "Earning T4 requires redesigning the governor, not retuning it." ### The Lean files, exactly The Lean 4 package sits in `kernel/aether/aether-verified/lean`, on toolchain `v4.7.0` with mathlib `v4.7.0`. Here is what it is, counted at commit `36b1400`: - 11 `.lean` files, of which `lakefile.lean` is the build file and four (`TestNat.lean`, `test_nat_chain.lean`, `test_nat_ineq.lean`, `test_sub_le.lean`) sit outside the `lake` roots. Two of those use names that no longer exist, so the build does not compile them. - 28 declarations in the five files the build compiles: `EpsilonTheorems.lean` 16 theorems and 1 private lemma, `Governor.lean` 3, `Betti.lean` 3, `Chebyshev.lean` 2, `Pruning.lean` 3. - `sorry` appears twice, both inside comments (`EpsilonTheorems.lean:7`, `lakefile.lean:15`). In code: 0. - Four statements conclude `True := trivial`. They are placeholders and prove nothing: `tss_separation_guarantee`, `gmc_entropy_nonincreasing`, `phkp_perfect_locality`, and `upper_bound_sound` in `Pruning.lean`. The repo's sentence on scope: "The Lean files prove algebraic side lemmas; no Lean statement is connected to kernel code." A boot `CERTIFIED` for T1 or T2 is the Rust kernel evaluating a hypothesis at its runtime parameters, not a Lean proof reaching into the kernel. The statements that matter, verbatim, with what each guarantees: **Lean 4 theorem tss_packing_bound** (0 sorry), [kernel/aether/aether-verified/lean/EpsilonTheorems.lean:51 @ 36b1400](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/aether/aether-verified/lean/EpsilonTheorems.lean#L51) ```lean theorem tss_packing_bound (L : ℕ) (θ_min : ℝ) (h_θ_pos : 0 < θ_min) (h_θ_lt_pi : θ_min < Real.pi) (cap_arg : (L : ℝ) * (Real.sin (θ_min / 2)) ^ 2 ≤ 4) : (L : ℝ) ≤ P_max θ_min := by unfold P_max have hθ2_pos : 0 < θ_min / 2 := by linarith have hθ2_lt_pi : θ_min / 2 < Real.pi := by linarith [Real.pi_pos] have hsin_pos : 0 < Real.sin (θ_min / 2) := Real.sin_pos_of_pos_of_lt_pi hθ2_pos hθ2_lt_pi have hsin_sq_pos : 0 < (Real.sin (θ_min / 2)) ^ 2 := pow_pos hsin_pos 2 rw [le_div_iff hsin_sq_pos] exact cap_arg ``` This proves that the packing count $L$ is at most $4/\sin^2(\theta_{\min}/2)$, given the cap-area inequality as the hypothesis `cap_arg`. The geometric fact itself is assumed, not proved; the theorem is the division step. **Lean 4 theorem scm_contraction** (0 sorry), [kernel/aether/aether-verified/lean/EpsilonTheorems.lean:84 @ 36b1400](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/aether/aether-verified/lean/EpsilonTheorems.lean#L84) ```lean theorem scm_contraction (α_min α_max ε_min ε_max : ℝ) (h_α_pos : 0 < α_min) (h_α_lt_1 : α_min < 1) (h_α_ord : α_min ≤ α_max) : telemetry_lipschitz α_min < 1 := by unfold telemetry_lipschitz linarith ``` `telemetry_lipschitz α_min` is defined as `1 - α_min`, so this proves $1 - \alpha < 1$ for $\alpha \in (0,1)$. Nothing about any map's Lipschitz behaviour is proved. The operator the kernel runs, $T(S) = (1-\alpha)S + \alpha S_{\text{pred}}$, appears in the Rust, not here. **Lean 4 theorem agcr_gain_margin_stable** (0 sorry), [kernel/aether/aether-verified/lean/EpsilonTheorems.lean:126 @ 36b1400](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/aether/aether-verified/lean/EpsilonTheorems.lean#L126) ```lean theorem agcr_gain_margin_stable (α β dt : ℝ) (h_α_pos : 0 < α) (h_β_pos : 0 < β) (h_dt_pos : 0 < dt) (h_margin : α + β / dt < 1) : 0 < governor_contraction_rate α β dt ∧ governor_contraction_rate α β dt < 1 := by unfold governor_contraction_rate have hβdt_pos : 0 < β / dt := div_pos h_β_pos h_dt_pos have hden_pos : 0 < 1 + β / dt := by linarith refine ⟨?_, ?_⟩ · -- 0 < 1 − α/(1 + β/dt) -- α < 1 + β/dt is implied by h_margin (α + β/dt < 1 < 1 + β/dt? not quite), -- but h_margin gives α < 1 - β/dt < 1 < 1 + β/dt have hα_lt : α < 1 + β / dt := by linarith have h_frac_lt_1 : α / (1 + β / dt) < 1 := by rw [div_lt_one hden_pos] exact hα_lt linarith · -- 1 − α/(1 + β/dt) < 1 have h_frac_pos : 0 < α / (1 + β / dt) := div_pos h_α_pos hden_pos linarith ``` This is the implication in equation (9): if the margin holds, $\rho$ lies in $(0, 1)$. It is not stated about the update map (8), so it does not say the governor converges, and at the runtime step its hypothesis is false (5.01). **Lean 4 theorem lyapunov_descent** (0 sorry), [kernel/aether/aether-verified/lean/AetherVerified/Governor.lean:33 @ 36b1400](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/aether/aether-verified/lean/AetherVerified/Governor.lean#L33) ```lean /-- **Scalar Lyapunov descent for a contraction.** If `e' = ρ · e` with `|ρ| ≤ 1`, then `V(e') ≤ V(e)`. Applies to a map of that form only; the Rust PD step is not one (see the module comment). -/ theorem lyapunov_descent (ρ e : ℝ) (h_ρ : |ρ| ≤ 1) : V (ρ * e) ≤ V e := by unfold V have hρsq : ρ ^ 2 ≤ 1 := by have := sq_abs ρ have hρ2 : |ρ| ^ 2 ≤ (1 : ℝ) ^ 2 := pow_le_pow_left (abs_nonneg ρ) h_ρ 2 simpa [sq_abs] using hρ2 have he2 : 0 ≤ e ^ 2 := sq_nonneg _ calc (ρ * e) ^ 2 = ρ ^ 2 * e ^ 2 := by ring _ ≤ 1 * e ^ 2 := by exact mul_le_mul_of_nonneg_right hρsq he2 _ = e ^ 2 := by ring ``` For a map $e \mapsto \rho e$ with $|\rho| \le 1$, $e^2$ does not increase. The file's own module comment says "No theorem here is about the Rust `governor_step`", and gives the counterexample at $\alpha = 0.01$, $\beta = 0.05$, $\Delta t = 1$, where the refined gain-margin condition holds: from $\varepsilon = 0.28$, $e_{\text{prev}} = 0.5$, $\delta = 0.2$ the step raises $|e|$ from 0.414286 to 0.414650. **Lean 4 theorem chebyshev_one_sided_sq** (0 sorry), [kernel/aether/aether-verified/lean/AetherVerified/Chebyshev.lean:67 @ 36b1400](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/aether/aether-verified/lean/AetherVerified/Chebyshev.lean#L67) ```lean theorem chebyshev_one_sided_sq {n : ℕ} (x : Fin n → ℝ) (μ : ℝ) (k : ℝ) (h_k : 0 < k) (σ_sq : ℝ) (h_σ_pos : 0 < σ_sq) (h_σ_def : σ_sq * (n : ℝ) = ∑ i, (x i - μ) ^ 2) : (((Finset.univ.filter (fun i => (k * k) * σ_sq ≤ (x i - μ) ^ 2)).card : ℝ)) * ((k * k) * σ_sq) ≤ σ_sq * (n : ℝ) := by have h_t_pos : 0 < (k * k) * σ_sq := by positivity have base := markov_count_bound (fun i => (x i - μ) ^ 2) (fun i => sq_nonneg _) ((k * k) * σ_sq) h_t_pos rw [h_σ_def] exact base ``` The number of points at or beyond $k$ standard deviations, times $k^2\sigma^2$, is at most $n\sigma^2$, which gives the ceiling $n/k^2$ that a memory-pruning rule ported from sigmoid uses. The link to the Rust is only that the Rust checks the inequality this file proves. **Lean 4 theorem tss_separation_guarantee** (0 sorry), [kernel/aether/aether-verified/lean/EpsilonTheorems.lean:67 @ 36b1400](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/aether/aether-verified/lean/EpsilonTheorems.lean#L67) ```lean /-- TSS Separation: Statement-level placeholder; great-circle distance on `S²` requires geometry we do not import here. -/ theorem tss_separation_guarantee {P : ℕ} (centroids : Fin P → ℝ × ℝ) (ε_adaptive : ℝ) (h_ε_pos : 0 < ε_adaptive) (h_ε_lt_2 : ε_adaptive < 2) (h_distinct : ∀ i j, i ≠ j → centroids i ≠ centroids j) : -- Statement-level placeholder until S^2 great-circle distance is imported. True := trivial ``` One placeholder, so the shape is visible: it compiles, carries a "0 sorry" chip, and guarantees nothing. The repo's hygiene gate strips comments, rejects `sorry`, `admit` and `axiom`, and accepts `True := trivial` only next to a `placeholder`, `skeleton` or `deferred` marker, which is how these four are allowed to stand. ### Paging and W^X Everything above runs in ring 0, so the page tables keep kernel code from being written. `memory/virt.rs` maps the image from its own PE section table: code is read-execute and never writable, read-only data is non-executable, and kernel data is writable and non-executable. `EFER.NXE` and `CR0.WP` are set. A walk splits a 48-bit virtual address into four 9-bit table indices and a 12-bit offset, as `translate_in_pml4` does with shifts and masks: $$ i_{\text{PML4}} = \left\lfloor \tfrac{VA}{2^{39}} \right\rfloor \bmod 2^9,\quad i_{\text{PDPT}} = \left\lfloor \tfrac{VA}{2^{30}} \right\rfloor \bmod 2^9,\quad i_{\text{PD}} = \left\lfloor \tfrac{VA}{2^{21}} \right\rfloor \bmod 2^9,\quad i_{\text{PT}} = \left\lfloor \tfrac{VA}{2^{12}} \right\rfloor \bmod 2^9,\quad \text{off} = VA \bmod 2^{12}. $$ **Figure 7.** A four-level x86_64 page walk through a model of the mapping memory/virt.rs builds: a 16 GiB identity map of 2 MiB read-write no-execute leaves, the image in 4 KiB pages with per-section flags, and a higher-half alias. Type a virtual address and the walk descends the four stacked tables, splitting the address into four 9-bit indices and a 12-bit offset. A switch flips to the RWX fallback the code takes when the PE section table is unreadable, and a bar panel counts the model's leaves by permission class, where a present, writable and executable leaf is the W^X violation the kernel's probe counts. Image sizes in the model are illustrative; the repo's measured scan of 24,004 pages is quoted beside the model's count, not produced by it. Colour key: structure: table pointer entries; proof: read-execute code leaves; baseline: read-only and read-write no-execute leaves; withdrawn: W+X leaf: a violation. **Measured: 0 / 24,004** W^X violations among scanned kernel-root pages. Control: before enforcement: every scanned kernel-alias page was writable and executable (4,311 of 4,311 in the results table; the repo's prose elsewhere says 4,310 of 4,310). Source: commit b3cf934, local QEMU, stated in the commit; chase_boot.sh wx. ## What's new in it Each pillar replaces a familiar approach, named here. **Accounting by quantity versus a model of shape.** Linux tracks resident pages, CPU time and queue depth per process. Seal OS asks the trainer for two scalars per step and keeps a fixed-size geometric summary of the run, 4,792 bytes per stream, from which it names the regime, without seeing the model. **Patience and gap thresholds versus revisitation.** The usual way to notice overfitting is a rule on the validation curve: stop after a number of epochs without improvement ([Keras `EarlyStopping`](https://keras.io/api/callbacks/early_stopping/) is the common form), or flag a gap between training and validation loss. `stratum` asks a different question: whether the validation curve comes back through values it has already visited. That is a property of the curve's shape in delay coordinates, which a constant offset does not change. The repo's negative control is a healthy run with exactly such an offset: the gap threshold flags it and `stratum` does not. **A KV cache in user space versus a prefix tree in kernel memory.** [vLLM's PagedAttention](https://arxiv.org/abs/2309.06180) manages KV blocks in user space against real models; `foliation` tests placement and constraint instead: the block table is a path in a tree the kernel owns, each block is a frame from the kernel's allocator, and eviction is restricted to free faces. The ranking it adds is measured against a null that removes the entrant count: on the boot trace, dropping the entrant count halves the hit rate (952 to 476 bp). The source calls the entrant count "a persistence proxy, and honestly also a frequency counter." The ranking is a lexicographic minimum over the free faces: $$ r(\ell) = \bigl(n_\ell,\; -d_\ell,\; u_\ell\bigr), $$ with $n_\ell$ the entrant count, $d_\ell$ the depth and $u_\ell$ the last-use tick. The locality-only null is $(-d_\ell, u_\ell)$ and LRU is $(u_\ell)$. In code: ```rust // kernel/seal-os/src/ml_engine/foliation.rs:225-233 @ 36b1400 fn rank_of(policy: Policy, l: &Leaf) -> [u64; 3] { match policy { Policy::Foliation => [l.entrants as u64, u64::MAX - l.depth as u64, l.last_use], Policy::Lru => [l.last_use, 0, 0], // from https://github.com/triton-lang/kernels/pull/22: sink + local window, as a null Policy::Locality => [u64::MAX - l.depth as u64, l.last_use, 0], Policy::Random | Policy::Belady | Policy::Adaptive => [0, 0, 0], } } ``` **Figure 8.** The kernel's KV prefix tree replayed on three synthetic traces (boot, chat with 4 live conversations, chat with 16 live), one descent at a time. Each square is one 4 KiB block; fill encodes how many distinct sequences ever entered it, a dashed outline marks a free face (resident, unreferenced, no resident children), a thick outline a pinned block, and a cross marks the victim chosen at that step. Five policies run on the same candidate set (foliation, LRU, the locality-only null, an adaptive duel, the Belady oracle) plus a seeded random null, with cumulative hit rate, end-of-trace bars and a pool-size sweep. On the boot trace foliation matches Belady and LRU scores 0; on the four-live chat trace foliation loses to LRU and to the locality null. Two invariants are counted live and must stay at zero: evictions of a referenced block and breaks of the rooted-subtree property. Colour key: parameter: entrant count of a block (colour bar); structure: selected policy's hit rate; the descending path; baseline: LRU; measured: the repo's QEMU value at pool 24; proof: live invariant at zero; self-check matches the repo's table; withdrawn: live invariant broken, or a self-check mismatch. **Comparing floats versus enclosing them.** Every discrete answer in the usual numeric pipeline (a component count from a distance threshold, a top-k from sorted scores, a nearest centroid from a dot product) is a bare comparison of computed floats. Seal OS replaces each bare comparison with an enclosure and a refusal. The cost of that change is concrete and the repo records it: the nearest-centroid index's reported hit rate falls from 1.0000 to 0.9818, because the 91 queries it could not certify go to a full scan instead of being counted as hits. **A theorem banner versus a refusing gate.** A system with proofs attached usually prints that they hold. Seal OS evaluates each hypothesis at its running parameters, prints which fail or were never checked, and its CI gate rejects a log that says otherwise. ## What no one else built For each piece I name the closest prior work and say what differs, including where the prior work already does it. **Verified kernels.** [seL4](https://sel4.systems) has a machine-checked proof of functional correctness of its C implementation in Isabelle/HOL ([Klein et al., SOSP 2009](https://dl.acm.org/doi/10.1145/1629575.1629596)). Seal OS has nothing comparable: its Lean files prove algebraic side lemmas and none is connected to kernel code. What Seal OS does instead is evaluate its theorem hypotheses at boot against the live parameters and refuse the ones that fail, including its own convergence claim. That is runtime checking of conditions, not verification, and I would not put it in the same category as seL4's proof. **Rust research kernels.** [Theseus](https://github.com/theseus-os/Theseus) (Boos et al., OSDI 2020) restructures OS state into runtime-composable cells in one address space; [Redox](https://www.redox-os.org) is a Rust microkernel with drivers in user space; [Asterinas](https://github.com/asterinas/asterinas) is a Rust framekernel that runs Linux binaries with `unsafe` confined to a small framework. All three are about structure, isolation or ABI, and on those axes Seal OS, a monolithic image with its own 69-call ABI and no ring-3 execution yet, is behind each of them. The difference is elsewhere: where its state lives (points on $S^2$ owned by Voronoi cells) and what it answers about the workload. **Learned and ML-aware OS components.** [LinnOS](https://www.usenix.org/conference/osdi20/presentation/hao) (Hao et al., OSDI 2020) runs a small neural network in the kernel to predict, per I/O, whether an SSD will be fast or slow, and revokes a request predicted slow so the application can fail over ([paper](https://www.usenix.org/system/files/osdi20-hao.pdf)). LinnOS trains on I/O traces collected from the workload and reports 87 to 97 % inference accuracy. `stratum` is the opposite construction: it learns nothing. It applies a fixed geometric rule, with a constant derived rather than trained, to two scalars the workload hands it. **Loops in delay embeddings.** The idea that a loop in a sliding-window embedding of a time series is a signal is not mine. [Perea and Harer's SW1PerS](https://arxiv.org/abs/1307.6188) scores periodicity by the maximum 1-dimensional persistence of a sliding-window point cloud. `stratum` uses the same family of idea for a different question (did the validation curve revisit its own values?) and in a narrower form: a single scale $\kappa\varepsilon^\ast$ instead of a persistence diagram, a cycle rank that upper-bounds $\beta_1$ instead of computing it, an exact monotonicity certificate for the zero case, and the band $\sqrt{8/3} < \kappa < \sqrt{3}$ derived from the geometry of a symmetric V and a monotone stretch. The derivation of that band, and the use of a fold score as a kernel-side training-regime signal, are the parts I did not find in the prior work I compared against. **Certified β₀.** A persistence library such as [Ripser](https://github.com/Ripser/ripser) computes the barcode and leaves the threshold to the user. Equation (2) says something equivalent in persistence terms: the count at $s$ is certified exactly when the band around $s$ falls in a gap of the 0-dimensional barcode. The mathematics is standard. What differs is the interface: a count is returned only with that gap as its certificate, and otherwise a refusal names the tree edge in the band, with a scale-safe distance so the heights themselves do not overflow. **Error-bounded filters.** The closest prior work to certify-or-refuse is [Shewchuk's adaptive-precision geometric predicates](https://people.eecs.berkeley.edu/~jrs/papers/robust-predicates.abstract) (1996/1997), and the filtered predicates in computational-geometry libraries that follow them. A predicate evaluates in floating point with an error bound; if the bound cannot certify the sign, it escalates precision until it can. Seal OS uses the same filter idea with [Higham's bound](https://epubs.siam.org/doi/book/10.1137/1.9780898718027) on three discrete answers (a component count at a scale, an attention top-k set, a nearest centroid on $S^2$), and on failure it does not escalate. It refuses, widens the row, or scans every centroid, and the refusal names the input. The filter is Shewchuk's idea; the refusal as a returned value across a kernel interface is the difference. **Prefix caching.** Here the prior work already does most of it. [SGLang's RadixAttention](https://arxiv.org/abs/2312.07104) keeps the KV cache in a radix tree and evicts least-recently-used leaves first, recursively, while protecting nodes the running batch uses. That is free-face eviction under LRU. [vLLM's automatic prefix caching](https://docs.vllm.ai/en/v0.9.0/design/automatic_prefix_caching.html) hashes each block together with the tokens of its prefix, evicts unreferenced blocks by LRU, and on a tie evicts the block at the end of the longest prefix first, which is a depth term. So the connected-subtree constraint is not new, and the syscall-facing cache in Seal OS, which defaults to LRU, is the RadixAttention policy run in kernel memory. What differs: the tree, its frames and its eviction live in the kernel and draw on the kernel's own allocator; a child is shared only on an exact token match even though the key is a 64-bit digest; and the entrant-count ranking is tested against a null, [Belady's oracle](https://doi.org/10.1147/sj.52.0078) and 32 random seeds on a trace where it wins and one where it loses, with both outcomes published. Whether any of this beats user-space placement has not been measured. **What survives the comparisons.** The mechanism I did not find elsewhere is the combination: one certify-or-refuse rule applied across component counts, attention top-k, nearest-centroid lookup on $S^2$, the zero case and the trend signal of a fold score with a derived scale band, and the kernel's own boot theorems, with the theorem refusal enforced in CI. It is not applied everywhere: a window that turns is scored from its Rips complex without a certificate, and ManifoldFS places files with an index that carries none. Each part has a named ancestor above; the combination, and the kernel refusing its own convergence theorem at the step it runs, are what I add. **Figure 9.** Nearest-centroid lookup on the sphere with its certificate. Centroids sit on S² with their Voronoi cells tinted; a latitude-longitude grid hashes each centroid to a cell. For a query you move, the index searches the 3×3 block of grid cells around it (the whole cap near a pole), and the answer is certified only when the distance to the best centroid in the block, plus 1e-9, does not exceed the distance from the query to the edge of the block; otherwise every centroid is scanned. The geodesic circle of that radius is drawn around the query. A batch panel runs seeded queries for five centroid counts and compares the certified answers with the same lookup with the certificate switched off. The 1,215-of-5,000 figure from the repo is quoted beside it, not reproduced: the old code is not in the checkout. Colour key: structure: best centroid in the searched block; the block patch; proof: true nearest centroid; CERTIFIED; measured: the bound circle; baseline: block-only answer when it is wrong; parameter: query position. ## Limitations These are stated as the repo states them, and where the code at `36b1400` has moved past a document, I say which. **No user mode.** No instruction has executed in ring 3. `/bin/init` is absent and the kernel falls back to its in-kernel desktop; with a `/bin/init` present, boot stops after loading it. The repo's record, made at `9ebbe2e`, says `syscall_entry` ran on the user's stack with no `swapgs`. At `36b1400` it executes `swapgs` and loads the task's kernel stack; since nothing has run in ring 3, that path has not been exercised from user mode. The record also says no task is ever scheduled; at this commit the scheduler adopts the boot thread in `init`, and I have not run it, so I make no claim either way. Application-processor bring-up exists and is not called. **No real models.** Every `stratum` and `foliation` number comes from a synthetic fixture. The README: "It has never seen a real model, and its verdict is advisory: nothing enforces it yet." On the chat trace foliation loses to LRU (5,284 against 8,068 bp), to the locality-only null (6,818) and to all 32 random seeds; with 16 live conversations it reverses again in one mutation build. The reported results are at one pool size (24 blocks); the source also replays each trace at six sizes from 8 to 48 blocks, which the repo's one-pool-size limit does not mention. Whether kernel placement of the KV cache buys anything over user-space PagedAttention has not been measured. **The κ band covers one shape.** The band in (6) is proved for a symmetric V only. A V descending at 0.005 and climbing at 0.01 per step scores 0.016 at $\kappa = 1.72$ and 0.406 only at 1.8, above the ceiling, so no choice of $\kappa$ inside the band detects it. **T4 is refused and needs a redesign.** No step size earns it, because the loop's gain is about $10^6$ and (9) assumes one. System calls 100 and 102 still print `FAILED` for every theorem that is not certified, eight of the ten, where the boot log, the ManifoldFS status and the theorem viewer print `CERTIFIED`, `NOT CERTIFIED` or `NOT CHECKED`. Seven of the ten boot lines are not checked. **Lean is not connected to kernel code.** Eleven `.lean` files, 0 `sorry` in code, four `True := trivial` placeholders, and the proved statements are algebraic side lemmas. None is linked to the Rust by refinement. **CI is red.** The QEMU job of run 36165748105 fails at the language-hygiene gate, on `scripts/ci_parity.sh`, and the gates after it in that job did not run. The kernel's own unit tests are not run by CI: the crate is outside the Cargo workspace, and the Kernel Tests workflow last executed on 2026-08-11 (514 of 514). Two of the 25 boot milestones are string matches that prove little. **One machine, no comparison.** Hardware is one QEMU configuration (q35, OVMF, AHCI, 4 GiB, no NIC). The GPU path ran only on its CPU fallback. Every cycle count is QEMU TCG with an emulated TSC. No comparison against Linux or Ubuntu has been run. Nothing here is production-tested. **Security is measured, not complete.** W^X holds since `b3cf934`, but the probe classifies per mapping; SMEP and SMAP were not exercised because the CI CPU model lacks them; KASLR randomises mappings, not the image base. `manifold_acl::check_access` denies access Linux permits: a uid-1000 process cannot read a 0644 root-owned file. ManifoldFS places files with its own cell index, `fs/voronoi_cap.rs`, which carries no certificate. No ext4; TLS accepts Ed25519 certificates only; no socket system call exposes the network stack. **The docs lag the code.** The drift I found at `36b1400`: - README and THEOREMS.md quote `9 of 10 theorems VERIFIED; ... T1-T3, T5 ACTIVE`; the code prints `2 certified, 1 not certified (T4), 7 not checked`, and T3 and T5 are not checked. - THEOREMS.md cites line ranges for `init_theorems` and a `verify_topology_theorems` symbol that is no longer in `kernel/seal-os/src`, and its scheduler and ManifoldFS line references point elsewhere at HEAD. - RESULTS.md's syscall-entry limit predates the `swapgs` entry now in `userspace.rs`. - RESULTS.md says `FitAction` is "enforced nowhere"; at HEAD a threshold is published that only an unconstructed preset reads, so the effect is the same. - The older papers under `docs/research/` claim a 241 to 1500× KV-cache speedup over LRU and a 21.8M-token horizon. Neither figure appears in README, RESULTS or THEOREMS, and neither is a current result. The current foliation numbers are the ones above. - RESULTS.md's subsystem table sums to 92,037 of the 96,110 kernel lines it states (at `9ebbe2e`); 4,073 lines are not itemised. **What failed.** - Withdrawn: T4 certified at boot Killed by: evaluated at Δt = 1, a step no runtime caller uses; at Δt = 0.01 the margin is 5.01 and the gate refuses (commit 3c14df0). - Withdrawn: The Rips-cycle floor for κ is √5/√3 ≈ 1.291 Killed by: wrong on both counts; pairing at matched value is not the closest approach, and the floor is √(8/3) (trajectory_shape.rs:41-45). - Withdrawn: κ = 1.5, with a ceiling of 2.0 argued from a straight arc Killed by: 1.5 detected no clean V at all; the 2.0 ceiling fell to a monotone staircase that scored 0.969 at 1.5 before two-step chords were quotiented (trajectory_shape.rs:62-64). - Withdrawn: The foliation policy beats LRU Killed by: a pool sweep: at 32 plaques and above LRU reaches the same 9.52 % ceiling, and on a pure-recency workload foliation ties LRU exactly (RECORD.md:810, 821-824). - Withdrawn: A bounded-degree verdict for the sparse filtration Killed by: it measured undirected degree, not the bounded quantity; retracted one iteration later (RECORD.md:4238-4246). - Withdrawn: Occupancy-flow equilibrium gives a doubling bound on S² Killed by: S² is Ahlfors 2-regular with constant at most 25, so the conjecture says nothing (RECORD.md:4959-4965). - Withdrawn: The paper's T3 degree bound, as derived Killed by: an earlier revision derived it wrongly; it is now labelled an assumption (epsilon_hollow\.tex:259-268). ## Read more [View the project](https://github.com/teerthsharma/Epsilon-Hollow) · [Source on GitHub](https://github.com/teerthsharma/Epsilon-Hollow) The project page at [teerth.dev/Epsilon-Hollow](https://teerth.dev/Epsilon-Hollow/) draws the results live. The source is [github.com/teerthsharma/Epsilon-Hollow](https://github.com/teerthsharma/Epsilon-Hollow), and the short link is [teerth.dev/epsilon-hollow](https://teerth.dev/epsilon-hollow). Key files at commit `36b1400`: - [README.md](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/README.md): the idea and the four pillars. - [docs/RESULTS.md](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/docs/RESULTS.md): every number with its provenance, equations (1) to (9), the Limits list. - [kernel/seal-os/src/theorems.rs](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/seal-os/src/theorems.rs): the boot theorem lines. - [kernel/aether/aether-verified/lean/](https://github.com/teerthsharma/Epsilon-Hollow/tree/36b1400d8c744c11434925f82af6668b84b33187/kernel/aether/aether-verified/lean): the Lean 4 package. - [certified_betti.rs](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/epsilon/epsilon/crates/aether-core/src/certified_betti.rs), [attention.rs](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/epsilon/epsilon/crates/aether-core/src/attention.rs), [trajectory_shape.rs](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/epsilon/epsilon/crates/aether-core/src/trajectory_shape.rs) and [foliation.rs](https://github.com/teerthsharma/Epsilon-Hollow/blob/36b1400d8c744c11434925f82af6668b84b33187/kernel/seal-os/src/ml_engine/foliation.rs): the certified kernels and the KV prefix tree. Related essays on this site: [separatrix](/separatrix), which grew out of this work, and [Aether-Lang](/aether-lang), whose `no_std` runtime the kernel embeds. The upstream fixes above are at [openxla/xla#46539](https://github.com/openxla/xla/pull/46539) (closed, and landed on main through Google's internal import as [3d5df1d](https://github.com/openxla/xla/commit/3d5df1da699bfb63cbeedaa56f09885c7974b06e)) and [tensorflow/tensorflow#124410](https://github.com/tensorflow/tensorflow/pull/124410) (merged; short link [teerth.dev/tensorflow-124410](https://teerth.dev/tensorflow-124410)). **Figure 10.** Where Epsilon-Hollow's work went, drawn only from links the owner confirmed on 2026-10-11. The repository led to openxla/xla pull request 46539 (closed; the change landed on main as commit 3d5df1d) and to tensorflow/tensorflow pull request 124410 (merged), and separatrix came out of it. Arrows follow how the owner words each relation. Focus or hover a node to read the owner's statement and what the pull request changed; the chips filter the map to the upstream pull requests or to separatrix. A merged or landed pull request is not evidence for any claim on this site. Colour key: structure: Epsilon-Hollow; ink-3: arrow: a relation the owner confirmed (led to, came out of). --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # Aether-Lang > A research language where persistent homology is a primitive, so a loop can stop when the Betti numbers of its data stop changing. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/aether-lang - Repository: https://github.com/teerthsharma/Aether-Lang - Project site: https://teerth.dev/Aether-Lang/ - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "Aether-Lang", 2026, https://teerth.blog/aether-lang - Languages: Rust, Lean 4, Aether - Question: Can a loop end on a change in the shape of its data instead of on a small float? - Headline result: [1, 1, 0] Betti vector at which the README tour's seal loop stops, at radius 1: one piece, one loop (control: printed by the tour program; not compared against a tuned scalar stopping rule) ## What it is Iterative numerical procedures almost always stop on a scalar. A residual is watched until it falls below a tolerance, and then the loop ends. The rule works, and it has two failure modes that anyone who has trained a model has met. A loss that oscillates in its third decimal forces a patience parameter or a moving average, and both are further hyperparameters. And a scalar compresses the whole state into one number before it thresholds it, so two residual fields with identical norms but different structure look the same to it. A float also wobbles in its last digits for reasons that have nothing to do with the data, and a program that decides with floats cannot tell a real change from rounding. Aether is the research language I wrote to try a different rule. Persistent homology is built into it. `topology.ph` computes a persistence diagram exactly over $\mathbb{F}_2$, in homology dimensions 0, 1 and 2, from a Vietoris–Rips filtration or from a lazy witness complex. `topology.betti` reads the Betti numbers at a radius: the number of pieces, loops and enclosed voids. Those are integers, and the loop keyword `seal` (also spelled 🦭) can end when they stop changing. The README puts the idea in one line: "A loss that goes from 0.0341 to 0.0339 is noise, but a loop count that falls from 3 to 1 is an event." The mathematics sits in a `no_std` Rust core against `libm`. The same engine that answers `topology.ph` in a terminal builds for a Cortex-M3 (`thumbv7m-none-eabi`) and for bare `x86_64`, and links into a kernel that owns its own allocator and scheduler. A tree-walking interpreter is the reference implementation; a bytecode VM sits behind it on coverage. Around the persistence engine there is a certified library: a linking number, a top-$k$, an argmin, a threshold and more, each of which returns its answer with the bound that proves it, or a typed refusal. [topological-ml-toolkit](/topological-ml-toolkit) is built on top of Aether-Lang. The premise is narrow, and I want to state it before anything else: some loops should end when the shape of the data stops changing, not when a float becomes small. Whether that beats a well-tuned scalar criterion on real problems has not been measured. The repository calls it "the most important missing number", and this essay does not supply it. Here is the mechanism, in the five steps the README uses: 1. A time series is delay-embedded into a cloud of points: each point pairs a sample with the samples a fixed lag behind it. 2. Each point grows a ball. As the radius grows, touching points are joined and triangles fill in. This is the Vietoris–Rips filtration. 3. Features are born and die. Pieces merge; loops appear and are filled. Each lifetime is a bar. 4. The bars alive at a radius give integer Betti numbers. 5. A `seal until` loop reads those integers as its stopping condition. The README's tour program does exactly this to eighteen samples of a sine. The part that matters is four lines: ```text let r = 0~ 🦭 until topology.betti(shape, radius=r)[0] == 1 { r = r + 0.25~ } ``` It prints `[one piece at radius, 1, betti, [1, 1, 0]]`: at radius 1 the thirteen-point cloud is one piece with one loop. The program never asks whether the signal is periodic. It asks for the shape. The loop is born at radius 0.93 and filled at 2.05, and the seal loop stopped inside that bar. Note what this loop is: it stops on an ordinary Boolean condition over a Betti number, not on stability. The stability form comes later, and it has a weakness the figure below makes visible. **Figure 1.** The README tour, run live. Eighteen samples of sin(0.7 t), delay-embedded in three dimensions with lag 2, give a 13-point cloud. The Vietoris–Rips complex grows with the radius r (edges and triangles admitted when their diameter is at most r); the exact F2 column reduction returns the barcode; the strip under it reads beta_0 and beta_1 at r. The repository reports \[9, 0, 0] at r = 0.5, \[5, 0, 0] at 0.8, \[1, 1, 0] at 1.0 and \[1, 0, 0] at 2.1, with one H1 bar from 0.93 to 2.05. Two stopping rules run on the same cloud: the tour's condition (stop when beta_0 = 1) and stable (stop when a pass leaves the Betti vector unchanged). Stop radii for other step sizes are computed by this figure's port of the engine and are not printed by the repository; with step 0.25 the stable rule stops at r = 0.5 on \[9, 0, 0]. Stops are ringed: solid green for the tour's condition, dashed grey for stable. Homology dimension is shown by marker, label and luminance, not by hue. Colour key: structure: the point cloud and the simplicial complex at radius r; parameter: the radius r and the loop step h the reader moves; proof: a stop where the rule’s condition holds; ink-2: the stable rule’s stop (dashed ring); measured: Betti vectors and bar ends printed by the repository. ## What it can do ### Three ways a loop can end A seal loop has one syntax, `seal until expr { body }`, and three behaviours depending on what the condition is. The README lists them side by side: ```text 🦭 until n >= 3 { ... } // a condition 🦭 until convergence(1e-6) { ... } // a tolerance 🦭 until stable(topology.betti(d, radius=r)) { ... } // an invariant ``` A plain condition is evaluated before each pass, and the loop stops when it is true. `stable(expr)` evaluates the watched expression before each pass and stops at the first pass that left it unchanged, by exact structural equality over numbers, Booleans, strings, lists and records. `convergence(ε)` is the scalar rule, written so it can be compared with the other two. Newton's iteration for $\sqrt 2$ is the repository's example: ```text let x = 1~ let n = 0~ 🦭 until convergence(1e-6) { n = n + 1~ x = (x + 2 / x) / 2~ x~ } print([n, x])~ ``` It prints `[5, 1.414213562373095]`. The fourth pass moves $x$ by about $2.1\times10^{-6}$ and the fifth by about $1.6\times10^{-12}$. With $\varepsilon=0$ the loop waits for an exact repeat, which floating point reaches one ulp below the correctly rounded $\sqrt2$. The rule, writing $v_i$ for the body's value after pass $i$ and $d$ for the max norm, is $$ \text{stop after the first pass } i \ge 2 \text{ with } d(v_i, v_{i-1}) \le \varepsilon . $$ **Figure 2.** The three seal-loop forms on programs the repository prints. Top: until n >= 3 stops after three passes. Middle: Newton's iteration for the square root of 2 under convergence(eps); bars are the move between successive passes on a log scale, the horizontal line is the tolerance, and the ringed bar is the pass where the rule stops (pass 5 at eps = 1e-6, printing 1.414213562373095). A staircase beside it shows how the stop pass moves as eps is swept, which is the point: a tolerance is still a knob. Bottom: the watched value per pass for four until stable programs, with the first repeat ringed: three replay the repository's printed outputs (a division count, a coupled rollout state, a face count) and one is computed live, the Betti vector of the tour cloud at step 0.25, which stops at r = 0.5 on \[9, 0, 0]. The figure ranks none of the rules. Colour key: measured: the move between successive passes; baseline: the scalar tolerance line; proof: the pass where a rule’s stop condition is met; parameter: the tolerance eps the reader moves. The repository is blunt about the middle form: "A tolerance loop is a scalar stopping rule." Only `stable` over a topological value is a topological stop. The cleanest example in the tree watches a certified face count of a planar drawing while spokes are added to a square, one per pass (`examples/planar_faces.aegis`): ```text let growing = [[[0, 0], [2, 0]], [[2, 0], [2, 2]], [[2, 2], [0, 2]], [[0, 2], [0, 0]], [[0, 0], [1, 1]]]~ let k = 1~ 🦭 until stable(euler(growing).faces) { if k < 4 { growing.push(spokes[k])~ k = k + 1~ } } ``` The program's own comment says the face count goes 1, 2, 3, 4, 4, and the loop seals with all four spokes in and four faces. ### Proof, or a typed refusal The second thing the language does is decide things without trusting a float. "A confident wrong answer is worse than no answer." Every decision in the certified library comes with the bound that proves it, or with a refusal that names its reason. The README's list of refusals is short enough to quote whole: ```text linking refused: Intersecting { segment_a: 0, segment_b: 0 } certify refused: BoundaryUndetermined { .. straddling: [1, 2] } arrangement refused: EdgesCross { first: 0, second: 1, count: 1 } persistent homology failed: TooManySimplices { max: 4096 } ``` "None of these ever becomes a zero, a default, or a best guess." The persistence engine follows the same discipline: when a filtration would exceed its simplex budget, it refuses before building it rather than truncating it. A certified top-$k$ is the simplest case. Each score $D_i$ carries a forward error radius $R_i$, built from $\gamma_n = nu/(1-nu)$ with $u$ the unit roundoff, and the code refuses outright when $nu > 0.5$. A set $T$ of $k$ indices is returned only if $$ \max_{i\in T}\,(D_i + R_i) \;<\; \min_{j\notin T}\,(D_j - R_j), $$ and otherwise the answer is `BoundaryUndetermined`. The tour makes this a stopping rule. It asks whether 1.0 is below 1.5 when the score may be off by a radius, and halves the radius until the answer is proven: ```text let radius = 2~ 🦭 until certified_threshold([1.0], radius, 1.5)[0] == "below" { print(["radius", radius, "verdict", certified_threshold([1.0], radius, 1.5)[0]])~ radius = radius / 2~ } print(["proven at radius", radius])~ ``` It prints `undetermined` at radius 2, 1 and 0.5, then `[proven at radius, 0.25]`. At 0.5 the interval's upper end is exactly 1.5, which is not strictly below it. **Figure 3.** Certified decisions on the reader's scores. Mode A replays the tour's threshold loop: each score is drawn as the interval from D - R to D + R against the threshold line, and the verdict stays undetermined until the interval clears the line (radius 0.25). Mode B is a top-k decision: with scores \[0, 1, 2, 10], radii \[12, 0, 0, 0] and k = 2, comparing only the rank-2 and rank-3 intervals would certify a set, but the rule max over T of (D + R) = 12 is not below min outside T of (D - R) = 2, so the library refuses with BoundaryUndetermined and names the straddling indices. Certified means rounding did not choose the answer on the stored inputs; it says nothing about whether the inputs are right. Colour key: baseline: the float point estimate; proof: a certified verdict and the inequality that holds; withdrawn: a typed refusal, with its message; parameter: the radius the reader moves. The linking number works the same way, and it is the second program in the tour: a square and a hoop threaded through it. **Measured: 7.07e-14** error bound on the Hopf-link linking number, which is therefore reported as the integer 1 ('linked'). Control: a split pair gives zero_linking; a touching pair is refused with Intersecting, not given a number. Source: README.md:78; PAPER.md:1251 @ 6d8d004 (printed output of the tour and Hopf-link programs). ### What the engine has measured The persistence engine is exact and slow, and the repository measures how slow. This is `scale_probe` on a regular circle with no radius cap, single core, `--release`, Windows 11: | dim | $n$ | pairs | seconds | | --: | ----: | -----: | ------: | | 0 | 200 | 200 | 0.049 | | 0 | 1,000 | 1,000 | 5.781 | | 0 | 4,000 | 4,000 | 335.049 | | 1 | 60 | 1,771 | 0.117 | | 1 | 120 | 7,141 | 2.202 | | 1 | 200 | 19,901 | 20.728 | | 1 | 300 | 44,851 | 131.343 | | 2 | 30 | 4,090 | 0.100 | | 2 | 50 | 19,650 | 1.859 | | 2 | 70 | 54,810 | 15.338 | Every pair count equals its closed form ($n$, $\binom n2+1$, $\binom n3+n$). Against the simplex count $m$, the local exponent in each of the seven intervals lies between 1.41 and 1.56, so running time is close to $m^{1.5}$ in all three dimensions. These are single runs on one machine with no confidence intervals; the repository says to read them as orders of magnitude. One ratio is robust to that, because it compares identical assertions on the same machine in one session: **Measured: 26×** persistence test time after replacing a linear face scan with a BTreeMap face index: 29.07 s to 1.10 s. Control: the same persistence_scale test with the linear face scan, before commit 27d70fa. n = one run per side. Source: PAPER.md:1652-1665 @ 6d8d004; same machine, --release. ### Where it went upstream **Upstream: **[triton-lang/kernels #22](https://github.com/triton-lang/kernels/pull/22), Add topology-derived sparse attention kernel (merged 2026-07-28). 3.48× Triton sparse vs dense-CSR at seq 4096 (80.9% block reduction); 1.04× at seq 1024. RTX 4060 Laptop GPU, 50 timing rounds, as stated in the PR. Aether-Lang led to my triton-lang/kernels#22 kernel. The repository describes `crates/aether-core/src/scheduled.rs` as a Rust port of that merged kernel: the same CSR block schedule from sink blocks, a local window and the top-$k$ blocks by an $H_0$ merge height. The port reproduces the schedule and matches dense masked attention to within $10^{-12}$ over four schedules and four block geometries, and its 16-block configuration visits 56 of 136 causal blocks, a 58.8% cut. That cut is cost. Whether the blocks it keeps are the right ones is a different question, and the repository's answer is no: at sequence length 512 the topological selection captures 5.6% of the achievable gain over random at an equal budget, and with top-$k$ raised to 32 it recovers 0.7446 of the attention mass against random selection's 0.8643. The upstream timings are the PR's; the repository does not reproduce them. ## How it was made ### The object > **Definition: Vietoris–Rips filtration and Betti numbers at a radius.** > > For a finite point set $X$ with a dissimilarity $d$, a non-empty subset $\sigma\subseteq X$ is a simplex that enters at its diameter: > > $$ > \mathrm{VR}_r(X)=\{\sigma\subseteq X,\ \sigma\neq\varnothing : f(\sigma)\le r\},\qquad f(\sigma)=\max_{u,v\in\sigma} d(u,v). > $$ > > Each persistence pair $(b,d)$ is a class born at radius $b$ and killed at $d$, with $d=\infty$ for essential classes. The Betti number at a radius counts the bars alive there: $\beta_k(r)=\#\{(b,d)\in\mathrm{Dgm}_k : b\le r last_value = value, RuntimeFlow::Return(value) => return Ok(value), RuntimeFlow::Break => break, RuntimeFlow::Continue => continue, } } Ok(last_value) ``` (Note: `interpreter.rs:1256-1276` @ `6d8d004`. `evaluate_condition` (lines 1279-1284) returns the error `condition must be boolean` for any non-Boolean value.) Three things follow from these lines. The stable rule compares the watched value before a pass with its value before the previous pass, so it accepts the first repeat: a one-pass window. A plain condition must be a Boolean, or the program stops with an error. And when the `for` loop runs out of its 1,000 passes, the function falls through to `Ok(last_value)`: hitting the cap is not an error. ### The Lean tree The repository has a Lean 4 tree, `Aether/`, of 11,637 lines on toolchain `leanprover/lean4:4.31.0`. It holds 48 `theorem` declarations, 47 in `Core.lean` and one in `VM.lean`, and a repo-wide grep finds no `sorry`. Two facts come first. CI does not build it, so nothing in the repository shows that it still compiles. And it is not connected to the Rust: there is no extraction, refinement proof or correspondence argument between `Core.lean` and the crates. It is a second implementation of the language in Lean (lexer, parser, static checker, pipeline, VM, with big-step relations and a fuel-bounded executor), exercised mostly by `example` blocks. Forty-five of the 48 theorems have one shape. Here is one: **Lean 4 theorem evalExprWithFnsRel_binary_add_sound** (0 sorry), [Aether/Core.lean:1124 @ 6d8d004](https://github.com/teerthsharma/Aether-Lang/blob/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/Aether/Core.lean#L1124) ```lean theorem evalExprWithFnsRel_binary_add_sound : EvalExprWithFnsRel [] [] (Expr.binary (Expr.num 2) BinOp.add (Expr.num 5)) (Value.num 7) -> evalExprWithFns 2 [] [] (Expr.binary (Expr.num 2) BinOp.add (Expr.num 5)) = some (Value.num 7) := by intro _ native_decide ``` The statement reads as soundness: if the big-step relation says `2 + 5` evaluates to `7`, the executor agrees. The proof does not use that. `intro _` takes the relation hypothesis and discards it, and `native_decide` then evaluates the right-hand side by compiled computation. What the theorem establishes is the conclusion alone: the fuel-2 executor returns `7` on the single closed program `2 + 5`. It does not show that the relation holds for that program, it says nothing about any other program, and `native_decide` adds Lean's compiler to the trusted base. All 45 closed instances, including the nine seal-loop ones, begin with `intro _`. They are checks of specific runs, not general guarantees. The repository's own account calls them closed instances discharged by `native_decide`; this essay adds only that the named hypothesis is never used. The other three are general in their arguments. **Lean 4 theorem lookup_bind_same** (0 sorry), [Aether/Core.lean:1096 @ 6d8d004](https://github.com/teerthsharma/Aether-Lang/blob/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/Aether/Core.lean#L1096) ```lean theorem lookup_bind_same (env : Env) (name : Ident) (value : Value) : Env.lookup (Env.bind env name value) name = some value := by unfold Env.bind Env.lookup simp ``` For every environment, name and value, looking up a name just bound returns the bound value. That is a real universally quantified lemma about the Lean environment, proved by unfolding and simplification. It says nothing about the Rust interpreter's variable table. **Lean 4 theorem eval_bound_var** (0 sorry), [Aether/Core.lean:1101 @ 6d8d004](https://github.com/teerthsharma/Aether-Lang/blob/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/Aether/Core.lean#L1101) ```lean theorem eval_bound_var (env : Env) (name : Ident) (value : Value) : evalExpr (Env.bind env name value) (Expr.var name) = some value := by unfold evalExpr exact lookup_bind_same env name value ``` The Lean evaluator, on a variable that was just bound, returns its value, for all arguments. It covers one expression form of the Lean evaluator and nothing else. **Lean 4 theorem compileCheckedFrameProgram_static_ok** (0 sorry), [Aether/VM.lean:804 @ 6d8d004](https://github.com/teerthsharma/Aether-Lang/blob/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/Aether/VM.lean#L804) ```lean theorem compileCheckedFrameProgram_static_ok {stmts : List Stmt} {result : SlotEnv × FrameFnEnv × List FrameOp} (h : compileCheckedFrameProgram stmts = some result) : ∃ checked, Static.checkProgramDetailed stmts = Except.ok checked := by unfold compileCheckedFrameProgram at h cases hs : Static.checkProgramDetailed stmts with | error found => simp [hs] at h | ok checked => exact ⟨checked, rfl⟩ ``` For every program, if the checked frame compiler succeeds, the Lean static checker accepted the program. The compiler is defined as a match on the checker's result, so this is close to its definition, and it is the only general theorem touching the VM. It does not say the compiled frame code is correct, or that the checker rejects anything bad. There is also a mismatch between the two implementations that matters for seal loops. The Lean semantics decides a loop condition with `truthy`: **Lean 4 def truthy** (0 sorry), [Aether/Core.lean:203 @ 6d8d004](https://github.com/teerthsharma/Aether-Lang/blob/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/Aether/Core.lean#L203) ```lean def truthy : Value -> Bool | Value.bool b => b | Value.num n => n != 0 | Value.float intPart fracMicros => intPart != 0 || fracMicros != 0 | Value.str value => value != "" | Value.list values => values != [] | Value.unit => false ``` and the seal rules use it (`truthy value = true -> StepStmt env (Stmt.seal (some condition) body) env (Flow.value Value.unit)`, `Core.lean:770-778`). In Lean, `seal until 1 { ... }` ends at once because `1` is truthy. The Rust interpreter rejects the same condition with `condition must be boolean`. The Lean seal rule also has no iteration cap, while the Rust loop stops at 1,000 passes. So even if the Lean tree were built and linked, it would describe a different language on these two points. ## What's new in it The usual way to stop an iterative procedure is a scalar tolerance, and the usual way to make it robust is to add patience. Keras's [`EarlyStopping`](https://keras.io/api/callbacks/early_stopping/) is the standard form: watch one monitored number, ignore changes smaller than `min_delta`, and stop after `patience` epochs without improvement. One continuous knob and one integer knob, both on a single compressed number. Aether moves the watched quantity from a float to a vector of integers computed from the data's shape, and moves the knob from a continuous threshold to a discrete window. The paper is careful about what that buys: "Topological convergence does not remove tuning. It moves it." An integer window is easier to reason about than a tolerance that interacts with the scale of a loss, and that is the claim, nothing larger. Today the window is one pass. The usual place for persistent homology is after the fact. A program computes a point cloud, hands it to [ripser](https://github.com/Ripser/ripser) or [GUDHI](https://gudhi.inria.fr/) from a notebook, and a person reads the barcode. The repository's comparison table marks ripser, GUDHI, giotto-tda and Dionysus as "Topology as control flow: no", which is a description of what they are for rather than a fault. In Aether the diagram is a value inside the running program, the Betti vector is something a loop condition can read, and the engine that produces it is the same `no_std` code that builds for a microcontroller and links into a kernel. The README's phrase for the goal is topology that makes decisions inside a running program, not topology that describes data afterwards in a notebook. The usual way a program decides "is this score below the threshold" or "are these curves linked" is a float comparison or a rounding. In Aether each such decision is a value that is either a certificate or a typed refusal, and refusals propagate as errors rather than as zeros. The persistence engine takes a budget in the same spirit: "The caps are a budget, not a correctness limit." A filtration that would exceed its simplex cap is refused before it is built, never truncated. The last difference is in how the repository treats its own claims. PAPER.md has a negative-results section, "what we got wrong", that lists withdrawn claims with the numbers that killed them, and a status ledger that marks the Lean tree as ungated rather than counting it. I list those in Limitations below. ## What no one else built Most of what is in Aether exists elsewhere, and better. The parts below are named against the closest prior work I found, with the concrete difference. **Persistence engines.** [Ripser](https://github.com/Ripser/ripser) ([Bauer 2021](https://doi.org/10.1007/s41468-021-00071-5)) computes Vietoris–Rips barcodes far faster than this engine, using cohomology, clearing and apparent pairs; the repository says plainly that ripser handles clouds of tens of thousands of points and this engine does not. [GUDHI](https://gudhi.inria.fr/), [Dionysus 2](https://mrzv.org/software/dionysus2/), [PHAT](https://research-explorer.ista.ac.at/record/1433) and [giotto-tda](https://giotto-ai.github.io/gtda-docs/) are libraries, in C++ with Python bindings or in Python, that return diagrams to a host program. Libraries in other general-purpose languages, such as Haskell's [Persistence](https://hackage.haskell.org/package/Persistence-2.0.1/src/README.md) and [TDA4j](https://scaladex.scala-lang.org/appliedtopology/tda4j), are the same shape. A [Chalmers paper](https://research.chalmers.se/publication/525991) implements persistent homology in the array language Futhark, but there the language is the vehicle for the algorithm, not a language with homology in its semantics. Aether's engine is the textbook reduction and claims nothing over any of these as an engine. **Stopping on topology.** This is not new as an idea. [Rieck et al. (ICLR 2019)](https://arxiv.org/abs/1812.09764) derive an early-stopping criterion for neural network training from neural persistence, a scalar summary of the persistence of the weight graph, used with a patience-style rule. [Melodia and Lenz](https://arxiv.org/abs/1911.02922) stop an iterative Voronoi interpolation when bottleneck and Wasserstein distances between diagrams cross a heuristic threshold. Both build a topological criterion into one algorithm, and both reduce it to a scalar compared against a threshold. **What is left, and what I claim.** In Aether the topological stop is not part of one algorithm. It is a loop form of the language: `seal until stable(expr)`, where `expr` can be the integer Betti vector produced by the language's own exact persistence engine, compared by exact equality, with no float in the stopping test. The same loop form takes a certified decision as its condition, as in the tour's threshold loop, so "stop when this is proven" is also a loop a program can write. And the engine that answers the condition is `no_std`. In the libraries, papers and languages above I found no loop construct whose termination condition is a Betti vector computed by a built-in persistence engine. That is a statement about my search, not a proof of absence, and the construct's value is exactly the unmeasured premise in the next section. **Not claimed.** The Lean tree is not a contribution of this kind. [CompCert](https://compcert.org/) and [CakeML](https://cakeml.org/) prove their implementations correct against their semantics; Aether's Lean tree is not connected to its Rust at all. The certified decisions are ports from my other projects (nerve, tangle, separatrix, sigmoid and others), so they are not new to Aether either. What Aether adds to them is that they are language values with typed refusals that a loop can wait on. ## Limitations **The core premise is unmeasured.** No controlled experiment compares stopping on Betti stability with stopping on a tuned scalar residual on real problems. The repository lists it first among its open items and calls it the most important missing number. Until it exists, nothing here says topological stopping is better. **No external parity.** The engine has not been checked against ripser, GUDHI, giotto-tda or Dionysus on shared fixtures. Every correctness statement in this essay is internal consistency: closed forms, negative controls, the stability bound, mutants. **`seal until stable` has a one-pass window.** Exact equality accepts the first repeat. The repository names this as the standing objection to topological stopping without a window. Figure 1 shows the effect on the tour's own cloud: with step 0.25 the stable rule stops at radius 0.5 on `[9, 0, 0]`, nine pieces, long before the loop that is the point of the example. That stop is computed by the figure's port of the engine, not printed by the repository; the window is the repository's own caveat. **A capped loop does not fail.** A seal loop that reaches its 1,000th pass returns its last value with no error. The README says loops are capped; it does not say the cap is silent. **Topology is opt-in, and one predicate is dead.** `convergence(ε)` is a scalar tolerance, `ConvergenceCond::BettiStable` is declared and never constructed, and the built-in `regress` statement stops on a sign pattern of its residuals, not on a filtration, with a tolerance of $10^{-6}$ whatever its `until:` field says. A seal loop is topological only when its author writes a topological condition. The paper's limits section still says the `convergence(ε)` spelling fails at runtime (PAPER.md:2141), while its §4.6 says it now runs and is pinned by tests (PAPER.md:1079-1081); one of the two passages is stale. **Scale.** $H_0$ at 4,000 points took 335 s and $H_1$ at 300 points 131 s. With every default in place, `topology.ph` refuses any cloud of 19 or more points, because the full 3-skeleton on $n$ points has more than 4,096 simplices from $n=19$: $$ m=\sum_{k\le K+1}\binom{n}{k+1},\qquad 18+153+816+3{,}060=4{,}047\ (n=18),\qquad 19+171+969+3{,}876=5{,}035\ (n=19). $$ **Figure 8.** The engine's measured ceiling, exactly as the repository reports it: seconds against the number of points n for homology dimensions 0, 1 and 2 (log-log), and the same times against the simplex count m, where the three lines nearly overlap with local slopes between 1.41 and 1.56. Every pair count is ticked against its closed form. A separate bar pair shows the face-index refactor, 29.07 s before and 1.10 s after. Single runs on one machine, no confidence intervals; nothing is extrapolated beyond the table and no other library is drawn. Colour key: measured: measured seconds; baseline: the linear face scan before the refactor; proof: a pair count equal to its closed form. The witness complex that `topology.ph` uses by default is not bounded as an approximation of the Rips diagram of the full cloud; no such bound is implemented or tested. **The kernel is not shown to boot.** It compiles for `x86_64-unknown-none`; there are no QEMU logs, and seven kernel tests never execute. **The Lean tree.** It is not built in CI and not linked to the Rust. Forty-five of its 48 theorems are closed instances whose proofs discard the relation they name. Its `truthy` accepts numbers and strings as loop conditions where the Rust interpreter rejects them, and its seal rule has no cap. **The headline stability test uses its own bottleneck.** The 12-case stability check compares diagrams with a bottleneck function defined in the test file. The crate's `diagram::bottleneck_distance` is cross-checked only in `diagram_distance.rs`. The paper describes the invariant check without that detail. **Evidence in the tree is not reproducible from the tree.** `crates/aether-core/test_results.txt` records a failed run (`manifold::tests::test_gatekeeper_branching`), and `output.txt` shows a path from an earlier name of the project. No mutation result file is tracked: `crates/aether-core/mutants.sh` defines a 26-mutant harness, and the paper reports 26 caught in the core and 26 in the GPU crate (the README's "52 of 52"), but no tally is in the repository. The README's 428 passing tests are a reported figure; I did not run the suites. The paper gives the Rust line count as 53,228; `wc -l` over the tracked `.rs` files under `crates/` at this commit gives 53,246 lines in 105 files. **Attention is synthetic, and the topological selection loses to random.** All attention results use synthetic keys. Placement, the share of the achievable gain over random that a selector captures at an equal budget, $$ \mathrm{pl}(S)=\frac{m(S)-m(R)}{m(O)-m(R)},\qquad m(S)=\frac1s\sum_{i=1}^{s}\sum_{j\in S_i}\mathrm{softmax}_i(j), $$ falls from 38.3% at sequence 128 to 5.6% at 512, and goes to −109% at top-$k$ 32. Block salience measures how isolated a block is, which is anti-correlated with attention mass by construction; selecting the lowest-salience blocks beats random. Every selector also allocates a dense `[seq, seq]` mask, so memory is quadratic. **Figure 9.** The repository's strongest negative result, rerun live. Tab 1: a selector that picks keys by Euclidean nearness to the query captures +0.884 of the achievable gain when key norms are equal and falls below random as norms spread (+0.202 at spread 2.0, -0.109 at 4.0, -0.285 at 8.0), reproducing the repository's table. Tab 2: the repository's recovered-mass table for the topological block selector across sequence lengths, plotted as data. Tab 3: single-linkage H0 structure of the keys and the routing gap ratio, showing that topological routing saves cost only when keys form clusters (cost 0.449 of dense on four clusters, 0.999 on uniform keys). Synthetic keys only; whether real attention keys carry H0 structure is unmeasured. Colour key: baseline: random selection, the floor; ink-3: oracle selection, the ceiling (placement 1); structure: the topological selector; withdrawn: placement below random; measured: the repository’s reported values; parameter: key-norm spread. **What failed.** - Withdrawn: Aether as the "Current World's Fastest Agentic AI Language" (the old PyPI description). Killed by: No benchmark supported it; removed (PAPER.md §8.1).. - Withdrawn: Euclidean proximity as a proxy for attention mass. Killed by: Placement +0.884 at zero key-norm spread, -0.285 at spread 8.0; the +0.884 is close to tautological under equal norms (PAPER.md §8.2).. - Withdrawn: An absolute radius of 0.6 as a key selector. Killed by: It selected 1.00 keys per row against 5.53 for the baselines; placements -3.573 to -4.177 were a budget artifact (PAPER.md §8.3).. - Withdrawn: Topological routing is sparse. Killed by: On uniform keys it costs 0.999 of dense; it pays only when keys have H0 structure (PAPER.md §8.4).. - Withdrawn: A budget-6 sliding-window fallback. Killed by: +0.014 placement on unstructured keys, which is random; the fallback is now dense (PAPER.md §8.5).. - Withdrawn: A placement of +7.614. Killed by: The ratio’s denominator was dominated by a dense fallback (PAPER.md §8.7).. - Withdrawn: seal until convergence(1e-6) stops on Betti stability. Killed by: It is a scalar max-norm tolerance (PAPER.md:2113).. - Withdrawn: ConvergenceCond::BettiStable is parsed and implemented. Killed by: It is never constructed (PAPER.md:2114).. - Withdrawn: The multiset of block saliences is the H0 barcode. Killed by: Centroids 0, 1, 10, 12 give saliences (9, 9, 2, 0) against deaths {1, 2, 9}; not invariant to block order either (PAPER.md:2118).. - Withdrawn: The governor update as first described. Killed by: It had the wrong sign and a non-convergent law (PAPER.md:2130).. ## Read more [View the project](https://github.com/teerthsharma/Aether-Lang) · [Source on GitHub](https://github.com/teerthsharma/Aether-Lang) - The project site, with every theorem and algorithm drawn: [teerth.dev/Aether-Lang](https://teerth.dev/Aether-Lang/). Short link: [teerth.dev/aether-lang](https://teerth.dev/aether-lang). - The repository: [github.com/teerthsharma/Aether-Lang](https://github.com/teerthsharma/Aether-Lang), read at commit [`6d8d004`](https://github.com/teerthsharma/Aether-Lang/tree/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50). - [README.md](https://github.com/teerthsharma/Aether-Lang/blob/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/README.md): the tour program and the three loop forms. - [PAPER.md](https://github.com/teerthsharma/Aether-Lang/blob/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/PAPER.md): the technical account, with 29 numbered results, the evaluation, "What we got wrong" and the limits. - The Lean tree, [`Aether/`](https://github.com/teerthsharma/Aether-Lang/tree/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/Aether), and its description in [docs/FORMAL_CORE.md](https://github.com/teerthsharma/Aether-Lang/blob/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/docs/FORMAL_CORE.md). - The engine, [`crates/aether-core/src/persistence.rs`](https://github.com/teerthsharma/Aether-Lang/blob/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/crates/aether-core/src/persistence.rs), and the seal loop, [`crates/aether-lang/src/interpreter.rs`](https://github.com/teerthsharma/Aether-Lang/blob/6d8d004717b154ad3f3b0bc2d9cd4d026e831a50/crates/aether-lang/src/interpreter.rs#L1202). - The upstream kernel: [triton-lang/kernels#22](https://github.com/triton-lang/kernels/pull/22) ([teerth.dev/triton-kernels-22](https://teerth.dev/triton-kernels-22)). - Related essays: [topological-ml-toolkit](/topological-ml-toolkit), which is built on Aether-Lang; [Epsilon-Hollow](/epsilon-hollow); [sigmoid](/sigmoid), the source of the coupling module. **Figure 10.** Where the certified library's modules came from, and the one upstream pull request this essay records. Left: the nine sibling repositories, each joined to the Aether module whose header names it as the port source (nerve to linking, separatrix to certify, sigmoid to coupling, and the rest). Right: triton-lang/kernels#22, merged 2026-07-28, joined to the scheduled module by an edge labelled, as the repository labels it, port of a merged PR. No edge says this project led to a pull request. A merged pull request records that a maintainer accepted a change; it is not evidence for any claim in this essay. Colour key: ink-3: sibling repositories and the edges that join them; structure: Aether modules; proof: the merged upstream pull request. --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # topological-ml-toolkit > Persistent homology as ordinary ML features, with a certificate that decides when one barcode can stand in for a whole moving point cloud. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/topological-ml-toolkit - Repository: https://github.com/teerthsharma/topological-ml-toolkit - Project site: https://teerth.dev/topological-ml-toolkit/ - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "topological-ml-toolkit", 2026, https://teerth.blog/topological-ml-toolkit - Languages: Rust, Python, C++, x86-64 assembly, CUDA - Question: When a point cloud moves, when can one persistence computation stand in for every frame? - Headline result: 1 of 3 persistence evaluations for a certified three-frame trajectory; every diagram equals dense recomputation within rtol 1e-10 (control: dense per-frame recomputation: 3 of 3) ## What it is topological-ml-toolkit is a library that turns the shape of a point cloud into numbers an ordinary estimator can use. It has two halves. `topoml-core` is a Rust crate that computes Vietoris–Rips persistent homology in dimensions 0, 1 and 2, with no `unsafe` code and a `no_std` build that needs only `alloc` and `libm`. `topoml` is a Python package that computes the same barcodes in a NumPy reference, turns them into Betti-curve and persistence-image features through transformers with scikit-learn's `fit`/`transform` shape, and carries a set of prototype diagnostics: metric covers and their nerves, Mapper graphs, sheaf residuals, a Scott-style fixed-point iteration, winding numbers, braid-crossing words and mesh Euler signatures. The question it answers is the one the docs open with: which clusters, loops and voids survive as the distance scale changes? A coordinate tells you where a point is. A barcode tells you that a cloud has one loop that is born at one radius and filled at another, and it says the same thing after you rotate, translate or re-embed the cloud. That matters to the three readers the design spec names: data scientists who know scikit-learn and PyTorch but not algebraic topology, systems programmers who want a bounded Rust core with optional native acceleration, and ML researchers who want to test whether a topology-derived feature or schedule carries any signal at all. The toolkit is built on top of Aether-Lang, my programming language that computes with shape. The only trace of that lineage inside this repository is one line in a planning document that asks for a taxonomy of topology families "not adequately covered by Aether/Epsilon". The object at the centre is the Vietoris–Rips complex. > **Definition: Vietoris–Rips complex.** > > For a finite point cloud $X$ with metric $d$ and a radius $\epsilon \ge 0$, a set of points is a simplex when every pair in it is within $\epsilon$: > > $$ > VR_\epsilon(X)=\{\sigma\subseteq X \;:\; \max_{u,v\in\sigma} d(u,v)\le\epsilon\}. > $$ > > Two points within $\epsilon$ give an edge, three pairwise-close points give a triangle, four give a tetrahedron. A simplex enters the filtration at its diameter, the largest pairwise distance inside it. As $\epsilon$ grows, simplices only arrive; none leave. Each homology class is born at one radius and dies at another, and the barcode in dimension $k$ is the list of those intervals: $$ B_k=\{(b_i,d_i)\}_{i=1}^{m}, \qquad p_i=d_i-b_i . $$ Here $b_i$ is the birth radius, $d_i$ the death radius and $p_i$ the lifetime. Classes that never die are essential; the code stores them with `death = None`. In the figure below you move the points yourself, and the barcode is computed by a port of the repo's own reducer. **Figure 1.** The Vietoris–Rips complex of a point cloud you can drag, grown by the radius slider ε. Edges appear when two points are within ε and triangles when all three pairs are; a loop's bar ends at the moment the last triangle filling it arrives. The barcode beside it is computed by a TypeScript port of the repo's Z/2 column reduction (python/topoml/core.py:200-286 @ 75fae45), and the finite H0 deaths are checked live against an independent minimum-spanning-tree computation. Presets include the repo's fixtures: three collinear points with H0 deaths 0.2 and 4.8, and the unit square whose H1 bar is alive at 1.1 and dead at 1.5. Colour key: structure: points, edges and triangles of the complex; seq-2: H0 bars (connected components) and the β0 curve; seq-3: H1 bars (loops) and the β1 curve; seq-4: H2 bars (voids), hatched, and the β2 curve; parameter: the radius ε and the simplices that just appeared at it; proof: the engine reproduces a value the repo asserts, or the H0 deaths match the independent minimum-spanning-tree check. ## What it can do There are no committed timing results in this repository. CI generates benchmark JSON and uploads it as a workflow artifact, and nothing of it is checked in. So everything I show in this section is a value the repo asserts in a test or in its end-to-end claim gate, or a count I derive from the code and label as derived. ### It agrees with ripser and GUDHI on the fixtures it checks The first job of a persistence library is to give the same barcode as the libraries people already trust. The repo has a CI job, `tda-baselines`, that installs ripser and GUDHI and runs both against the Python reference on two exact fixtures. **Measured: 0.2, 4.8** finite H0 deaths of the points (0,0), (0.2,0), (5,0), equal in topoml, ripser and GUDHI. Control: ripser.ripser and gudhi.RipsComplex on the same points, asserted equal. n = 3 points, max_dim 0, max_radius 10.0. Source: python/tests/test_tda_baseline_parity.py:38-60 @ 75fae45; CI job tda-baselines; not re-run for this essay. **Measured: β₁ = 1 → 0** unit-square loop alive at radius 1.1 and dead at 1.5, the same counts in topoml, ripser and GUDHI. Control: ripser and GUDHI intervals counted with the same half-open rule. n = 4 points, max_dim 1, max_radius 2.0. Source: python/tests/test_tda_baseline_parity.py:63-85 @ 75fae45; CI job tda-baselines. The deaths are easy to check by hand. The two close points merge at their distance, 0.2; the far point joins at $5-0.2=4.8$. The square's four sides arrive at 1, closing a loop; both diagonals and all four triangles arrive at $\sqrt2\approx1.414$ and fill it, so the loop is alive at 1.1 and gone by 1.5. The same three-point fixture is asserted in the Rust crate as well, as $\beta_0 = 3, 2, 1$ at radii 0.1, 0.3 and 6.0. ### It turns a barcode into a fixed-width row A barcode has a variable number of bars, and an estimator wants a fixed number of columns. The toolkit's main encoder samples Betti numbers at chosen radii. The Betti number at radius $\epsilon$ counts the bars alive there, with a half-open interval, so a bar is no longer counted at the radius where it dies: $$ \beta_k(\epsilon)=\big|\{\,i : b_i\le\epsilon` with a const-generic dimension; coordinates must be finite. The entry point validates its configuration before building anything: homology dimension at most 2, a radius that is neither negative nor NaN, at least one point and no more than `max_points`. Every simplex push checks `max_simplices`, which the API reference calls a "fail-fast cap during complex construction"; the same page calls `max_points` a "fail-fast cap before simplex explosion". The defaults are 64 points and 65,536 simplices. Those defaults have a concrete meaning, which I derive here from the enumeration (the repo does not state it). At infinite radius the complex for `max_dim = 1` on $n$ points has $n+\binom n2+\binom n3$ simplices; for $n=64$ that is $64+2{,}016+41{,}664=43{,}744$, inside the cap. For `max_dim = 2` the tetrahedra add $\binom n4$, and the largest $n$ that fits is 35 ($35+595+6{,}545+52{,}360=59{,}535$; $n=36$ gives 66,711). A finite radius admits more points, because simplices above it are never pushed. ```rust // crates/topoml-core/src/lib.rs:379-393 @ 75fae45 pub fn persistent_homology( cloud: &PointCloud, config: PersistenceConfig, ) -> Result { validate_config(cloud.len(), config)?; let source = match config.complex_kind { ComplexKind::VietorisRips => cloud.clone(), ComplexKind::Witness { max_landmarks } => select_landmarks(cloud, max_landmarks)?, }; validate_config(source.len(), config)?; reduce_z2( &build_vietoris_rips_simplices(&source, config)?, config.max_homology_dim, ) } ``` The `Witness` arm is worth reading closely. It selects up to `max_landmarks` points by farthest-point sampling from point 0 and then builds the same Vietoris–Rips complex on those landmarks. That is landmark subsampling followed by Vietoris–Rips; it is not a witness complex, in which the non-landmark points decide which landmark simplices exist. The Python package is a separate implementation, not a binding. It builds with setuptools, and no file under `python/` imports the Rust crate. Its reference enumerates every subset of size 2 to `max_dim + 2` with `itertools.combinations` and has no point or simplex cap. The two agree on the shared fixtures; the ripser and GUDHI parity tests run against the Python side. ## What's new in it The usual way to get persistence for a moving point cloud is to recompute it for every frame. The repo's own dense baseline does exactly that, one `persistent_homology` call per frame. The toolkit adds one case where that work is provably unnecessary, and a check that decides at run time whether the case holds. Suppose every pairwise distance in frame $t$ is the same multiple of the distance in frame 0: $$ d_t(x_i,x_j)=s_t\,d_0(x_i,x_j),\quad s_t>0 \;\;\Longrightarrow\;\; VR_r(X_t)\cong VR_{r/s_t}(X_0),\qquad [b,d)\mapsto[s_tb,\,s_td),\qquad \beta_k^{(t)}(r)=\beta_k^{(0)}(r/s_t). $$ Every simplex's diameter is multiplied by $s_t$, so the sorted order is unchanged and the reduction pairs the same simplices. The barcode of frame $t$ is the barcode of frame 0 with each finite endpoint multiplied by $s_t$; essential bars stay essential. Rotation and translation leave distances alone, so any similarity motion of the whole cloud qualifies. `persistence_similarity_trajectory` tests this condition numerically. Let $\delta_t(i,j)$ be the frame-$t$ distance of pair $i\theta$. The scale is the least-squares fit on $U$, and the distortion is the worst relative residual over all pairs: $$ s_t=\frac{\sum_{U}\delta_0\,\delta_t}{\sum_{U}\delta_0^{2}}, \qquad \operatorname{dist}_t=\max_{i0}\ \max\!\left(\frac{\lvert\rho_{\max}-s\rvert}{s},\ \frac{\lvert\rho_{\min}-s\rvert}{s}\right) \;=\;\frac{\rho_{\max}-\rho_{\min}}{\rho_{\max}+\rho_{\min}} . $$ Stretch a unit square by $a=0.7$ along $x$: the horizontal sides have ratio 0.7 and the vertical sides ratio 1, so the distortion is at least $0.3/1.7\approx0.18$, nine orders of magnitude above the tolerance. The repo's test uses the stretch $(0.7,\ 1.1)$ on twelve random points, together with a squaring map and a single moved point, and asserts that all three fall back to dense recomputation. **Figure 8.** A point cloud moving through T frames, drawn as a ghost trail. The certificate is the repo's: least-squares scale per frame, worst relative pair-distance residual, tolerance 1e-10 (python/topoml/persistence.py:61-111 @ 75fae45). In similarity mode (the benchmark's own trajectory: shrink by e^(−0.15u), rotate about z, translate) the certificate passes, persistence is computed once, and the rescaled bars sit on the dense bars for every frame. Switch to anisotropic, nonlinear or one-point motion and the distortion bars jump far above the tolerance and the mode becomes dense fallback; force reuse anyway to see the error it would have made. The evaluation counter and the work reduction 1 − evaluations/T are counts, not timings. The only timing is the “Time it here” button, which measures H0 dense evaluations against certificate plus one evaluation in this TypeScript port, best of 3, on the reader's device only; the repo stores no timings. Colour key: structure: reused bars, rescaled from frame 0; baseline: dense bars, recomputed for the frame; proof: frame certified, distortion at or below tolerance; reused bar ends that coincide with the dense bar; withdrawn: frame not certified, distortion above tolerance; parameter: the selected frame and the tolerance. The repo asserts this in three places, each against dense recomputation: **Measured: 1 of 3** persistence evaluations for a unit square scaled by 1, 0.75, 0.5 and translated by (i, −i); every reused diagram equals the dense one within rtol 1e-10, atol 1e-12. Control: dense per-frame persistence, 3 evaluations. n = 3 frames, 4 points, max_dim 1, tolerance 1e-10. Source: benchmarks/e2e_claims.py:61-88 @ 75fae45; E2E claim gate in CI. **Measured: 1, 0.8, 0.55** scales recovered by the certificate on a rotated, translated, scaled 12-point circle, to rtol 1e-10, with max distortion at or below 1e-10 and diagrams equal to dense. Control: the known scales used to build the trajectory, and dense per-frame persistence. n = 3 frames, 12 points, max_dim 1. Source: python/tests/test_persistence_similarity.py:65-87 @ 75fae45. **Measured: 2/3** work_reduction field in the similarity benchmark's JSON for 3 frames: 1 reuse evaluation against 3 dense. Control: dense_persistence_evaluations = 3. n = 8 points, 3 frames, 2 repetitions, 0 warmups (CLI smoke test). Source: python/tests/test_persistence_similarity.py:205-237 @ 75fae45. In general the count is $1-1/T$ for $T$ certified frames: 0.875 at the CI's $T=8$ and 0.9375 at $T=16$. Those two values are derived by me from the formula in `benchmark_persistence_similarity.py`; they are not stored results. The count also leaves out the certificate's own cost, $T\binom n2$ distances, which is small next to a reduction over $\binom n3$ triangles but is not zero. The benchmark itself does measure time, with randomised run order, every raw paired sample and a paired 99% bootstrap interval on the geometric-mean speedup, but its output lives only in CI artifacts, its stated scope is "Python reference H0 on exact synthetic similarity trajectories", and none of its numbers is in the tree. So I quote none. The other thing this project does differently is the claim discipline around all of this. The README puts it as "This repository treats performance statements as claims that must be executable." Every active capability in the previous section is an assertion in the end-to-end gate or a test, every backend has a list of gates that must pass before it is selected, and a page in the docs lists what is not claimed. This project documents, in the same place, what it has not shown. ## What no one else built The specific thing here is small and exact: a public function that takes a trajectory, decides from pair distances alone whether every frame is a uniform scaling of the first within a stated relative tolerance, and if so returns all diagrams from one persistence computation, and otherwise says why not and recomputes. Around it sit a Rust core that builds without the standard library and refuses inputs over its caps, and a Python layer that exposes the result as scikit-learn-shaped features. Each part has close prior work, and the differences are narrower than the word "new" suggests. **Vineyards.** Cohen-Steiner, Edelsbrunner and Morozov, [Vines and vineyards by updating persistence in linear time](https://doi.org/10.1145/1137856.1137877) (SoCG 2006), handle the general version of this problem: when a filtration changes, the pairing is updated in linear time per transposition of the simplex order, so any continuous motion can be tracked. [Dionysus](https://mrzv.org/software/dionysus2/) implements vineyards. Under exact uniform scaling the simplex order never changes, so a vineyard update would perform zero transpositions. The toolkit handles only that case, which is much narrower, and what it adds is the cheap test that recognises it from distances, the closed-form rescaling, and an explicit fallback when the test fails. It does not track motion in general. **Stability.** Cohen-Steiner, Edelsbrunner and Harer, [Stability of persistence diagrams](https://doi.org/10.1007/s00454-006-1276-5) (2007), bound how far a diagram can move when the filtration moves; that theorem is what gives the reuse bound above. A user could reuse a diagram for any nearly similar frame and accept an error up to that bound. The toolkit declines to: it reuses only at a relative tolerance of $10^{-10}$, so the reused diagram is the dense one to within floating-point noise, and anything looser goes to recomputation. **Vietoris–Rips engines.** [Ripser](https://github.com/Ripser/ripser) (Bauer, [J. Appl. Comput. Topol. 2021](https://doi.org/10.1007/s41468-021-00071-5)) computes Rips barcodes through persistent cohomology with clearing, implicit boundary matrices and apparent pairs; [Ripser++](https://github.com/simonzhang00/ripser-plusplus) moves that onto the GPU; [PHAT](https://github.com/blazs/phat) offers several boundary-matrix reduction strategies; [GUDHI](https://gudhi.inria.fr/) provides Rips, alpha and real witness complexes on a simplex tree. Each uses reduction machinery the toolkit lacks or covers more kinds of complex than the toolkit's reducer, which is the plain standard algorithm with a linear-scan face lookup. The toolkit claims no speed against any of them; it uses ripser and GUDHI as the reference its fixtures must match. **ML features.** [giotto-tda](https://github.com/giotto-ai/giotto-tda) ([JMLR 2021](https://jmlr.org/papers/v22/20-325.html)) already wraps Vietoris–Rips persistence, Betti curves, persistence images, Takens embedding with automatic parameter search, and Mapper as scikit-learn transformers; [persim](https://persim.scikit-tda.org/) in scikit-tda provides persistence images. Persistence images themselves are from Adams et al., [Persistence Images: A Stable Vector Representation of Persistent Homology](https://jmlr.org/papers/v18/16-337.html) (JMLR 2017), who integrate a weighted sum of Gaussians over each pixel. The toolkit's version evaluates the sum at grid nodes with weight equal to persistence. Its transformers are not a contribution over giotto-tda; they are the interface the rest of the library hangs on. **Rust.** [LoPHAT](https://github.com/tomchaplin/lophat) is a Rust library with Python bindings that reduces general filtered complexes in parallel, lock-free, with clearing. It takes a complex as input and depends on rayon. `topoml-core` builds the Vietoris–Rips complex itself, runs single-threaded, compiles without `std`, and returns an error rather than allocating past its caps. I did not find another Rips crate that states a `no_std` build, but I did not search exhaustively, and I make no claim that it is the only one. What survives that comparison is the combination: an exact-reuse path whose acceptance test is stated as a number, whose fallback is a documented mode rather than a silent approximation, and which ships beside a claim ledger that lists its own gaps. The ledger below runs the repo's computed claims again in your browser. **Figure 9.** The repo's active claims as a table, each row with its expected value and source at commit 75fae45. Two rows are recomputed in the browser by a short TypeScript port, the H0 cluster merges and the H0 deaths of the three-point fixture; the other eleven keep their expected value and source and are marked not run. Reproducing a row shows that the port agrees with the asserted value, not that the Python library does; the Python run is the repo's CI. Backends that need hardware (C++, AVX-512, CUDA, Triton, PyTorch, TensorFlow) are listed with their gates, and four claims the repo explicitly does not make are listed with them. Colour key: proof: computed here and equal to the asserted value; withdrawn: computed here and different from the asserted value; refused: gated or explicitly not claimed by the repo. ## Limitations The sparse-attention side is the clearest place where the repo's discipline produces a null result. Its `TritonScheduleBuilder` picks a key budget from sink tokens, a local window and farthest-point landmarks, and the repo's performance gate scores any sparse schedule by its error against dense attention: $$ \text{relative }L_2=\frac{\lVert y_{\text{dense}}-y_{\text{sparse}}\rVert_2}{\lVert y_{\text{dense}}\rVert_2}. $$ On the repo's own six-key fixture (keys at 0, 0.1, 2, 2.1, 5, 5.1 on a line, budget 4, one sink, window 1, two landmarks), the topology schedule selects keys $(0,3,4,5)$, and the local-window baseline selects the same $(0,3,4,5)$. The E2E gate asserts both. The schedule builder's own claim scope reads "schedule construction only; no Triton kernel or sparse-attention speedup claim", and the docs require dense SDPA or FlashAttention baselines before any speedup claim. **Figure 10.** The repo's topology-derived key schedule, a local-window schedule and a same-budget random schedule, side by side on the same keys, each scored by relative L2 and cosine similarity against dense softmax attention for one query. Keys and values are synthetic and seeded; the random schedule is summarised over 200 draws, where the repo uses one. On the repo fixture the topology and local schedules are the same set. The best state the gate can show is 'accuracy gate passed — necessary, not sufficient', because a promotion also needs schedule and kernel timings against dense attention, which this figure cannot provide and the repo does not have. Colour key: structure: keys selected by the topology schedule; baseline: keys selected by the local-only and random schedules, and the grey band of the 200 random draws; line-2: causal keys not selected (grey dots and bars, bar height = dense weight); keys after the query are outlined; parameter: the query key (diamond); proof: relative L2 within the chosen tolerance; withdrawn: relative L2 above the chosen tolerance; ink-2: the chosen relative-L2 tolerance (dashed). The rest, stated as the repo and its code show them: - **No measured results in the tree.** Every timing and every speedup interval is a CI artifact that is not committed. The parity tests cover two small fixtures and, in the repo's words, "do not claim runtime leadership, large-scale barcode equivalence, or C++/ASM/GPU acceleration". The E2E gate "does not claim a speedup over ripser, GUDHI, sklearn, dense SDPA, FlashAttention, or any framework kernel". - **Reuse is Python-only and H0 in its benchmark.** The certificate exists only in the Python package. The benchmark's scope is the Python reference on H0 for exact synthetic trajectories, and the docs say it does not compare against ripser, GUDHI, Rust, CUDA or Triton. The tolerance "is an implementation certificate, not proof that noisy measured motion follows an exact physical law"; measured trajectories with noise will almost always fall back. - **Accelerators cover sub-operations only.** C++ does H0; assembly does squared-L2 dispatch; CUDA does pairwise L2 and threshold edges; Triton does pairwise L2 and CPU-side schedule construction. H1/H2 native reduction, GPU persistent homology, topology-guided sparse attention and framework-native PH kernels are gated and unclaimed. Parity tests for the native and GPU paths skip on machines without a compiler or CUDA. - **Cost.** The Python reference enumerates every subset up to size `max_dim + 2` and has no cap. The Rust reducer finds each face by a linear scan over earlier simplices. Both use the plain standard reduction, which is cubic in the number of simplices in the worst case (a standard result, not a repo statement). The Rust circle test asserts only $\beta_1\ge1$ at radius 1.0. - **Coverage of the CI parity job.** The default `python` job installs only the test extra, which has neither ripser nor GUDHI; the parity tests run only in the separate `tda-baselines` job. - **Training baselines.** The forest's training score of 1.0 in the E2E gate is measured on the same four clouds it was fitted on. `topological_sample_weights` takes differences across the whole concatenated row, including across the $\beta_0$-to-$\beta_1$ block boundary. The braid word "is not a complete knot polynomial or link invariant". The Reeb relation is documented, but the only API is `mapper_graph`. - **Planned files that do not exist.** The backend plan names a Triton adapter module, a C++ header and README, an assembly README and two backend docs that are absent at this commit. There is no CHANGELOG. A set of places where the documentation or a test name says more than the code does: **What failed.** - Withdrawn: The landscape doc lists 'persistent homology over bounded Vietoris-Rips and witness complexes' (docs/concepts/topology-landscape.md:17). Killed by: ComplexKind::Witness selects farthest-point landmarks and builds Vietoris–Rips on them; no witness relation is computed, and no Rust test exercises the variant (crates/topoml-core/src/lib.rs:384-392, 414-449; tests/persistence_contract.rs:1-58).. - Withdrawn: Persistence landscapes, total persistence, persistence entropy and Chebyshev radii are given as feature formulas (docs/concepts/ph-feature-vectors.md:6-31, 58-81). Killed by: None is implemented in python/ or crates/; the doc itself says 'Direct landscape encoding is not implemented yet' (ph-feature-vectors.md:81).. - Withdrawn: A 'GPU backend semantic contract' test covers the CUDA and Triton kernels. Killed by: It runs no CUDA or Triton code: it compares NumPy re-implementations with hard-coded arrays and checks that symbol names appear in the source files (python/tests/test_gpu_backend_semantic_contract.py:11-24, 102-141).. - Withdrawn: mesh_euler_characteristic reports closed_orientable and a genus. Killed by: The flag only checks that every edge lies in exactly two faces; face orientations are never compared, so a closed non-orientable mesh passes (python/topoml/topology.py:550-561).. - Withdrawn: TopologyRandomForestClassifier applies topological sample weights. Killed by: It applies them twice: as bootstrap sampling probabilities and again as weights in each stump's loss (python/topoml/training.py:170-194).. ## Read more [View the project](https://github.com/teerthsharma/topological-ml-toolkit) · [Source on GitHub](https://github.com/teerthsharma/topological-ml-toolkit) The project's own site, built from the repo's MkDocs pages, teaches topology first, then ML usage, then the backend and benchmark contracts. The source is [teerthsharma/topological-ml-toolkit](https://github.com/teerthsharma/topological-ml-toolkit), and everything in this essay is read at commit [`75fae45`](https://github.com/teerthsharma/topological-ml-toolkit/tree/75fae45a1b0ea852c5a7c696d8227e12c7ba1333). Key files at that commit: - [README.md](https://github.com/teerthsharma/topological-ml-toolkit/blob/75fae45a1b0ea852c5a7c696d8227e12c7ba1333/README.md): the working surface, backend table and benchmark discipline. - [python/topoml/persistence.py](https://github.com/teerthsharma/topological-ml-toolkit/blob/75fae45a1b0ea852c5a7c696d8227e12c7ba1333/python/topoml/persistence.py): the similarity certificate and reuse path. - [python/topoml/core.py](https://github.com/teerthsharma/topological-ml-toolkit/blob/75fae45a1b0ea852c5a7c696d8227e12c7ba1333/python/topoml/core.py): the Python reference reducer. - [crates/topoml-core/src/lib.rs](https://github.com/teerthsharma/topological-ml-toolkit/blob/75fae45a1b0ea852c5a7c696d8227e12c7ba1333/crates/topoml-core/src/lib.rs): the `no_std` Rust core. - [benchmarks/e2e_claims.py](https://github.com/teerthsharma/topological-ml-toolkit/blob/75fae45a1b0ea852c5a7c696d8227e12c7ba1333/benchmarks/e2e_claims.py) and [docs/benchmarks/e2e-claims.md](https://github.com/teerthsharma/topological-ml-toolkit/blob/75fae45a1b0ea852c5a7c696d8227e12c7ba1333/docs/benchmarks/e2e-claims.md): the claim gate and what it does not claim. - [docs/concepts/persistent-homology.md](https://github.com/teerthsharma/topological-ml-toolkit/blob/75fae45a1b0ea852c5a7c696d8227e12c7ba1333/docs/concepts/persistent-homology.md): the maths, including exact reuse under similarity motion. Related essays on this blog: - [Aether-Lang](/aether-lang), the language and runtime this toolkit is built on, where an H0 barcode decides when topological routing of attention pays. - [nerve](/nerve), on global topology (knot type, linking) that local descriptors merge away. - [sigmoid](/sigmoid), which builds its H0 barcode from a minimum spanning tree and uses it inside a world model. - [planimeter](/planimeter), where single-linkage merge heights, the same numbers as the H0 deaths here, decide which endpoints are the same point. Upstream: [google-deepmind/mujoco#3396](https://github.com/google-deepmind/mujoco/pull/3396), "Remove quadratic scratch from native island discovery" (short link [teerth.dev/mujoco-3396](https://teerth.dev/mujoco-3396)), and the [MuJoCo 3.11.0 changelog](https://mujoco.readthedocs.io/en/3.11.0/changelog.html#version-3-11-0-july-27-2026) entry that credits it. --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # nerve > Can a local atomic descriptor see how polymer chains are linked and knotted? A Rust harness that counts what the descriptor merges. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/nerve - Repository: https://github.com/teerthsharma/nerve - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "nerve", 2026, https://teerth.blog/nerve - Languages: Rust - Question: Which chain topologies does a local atomic descriptor send to the same output, and can the cost of that be counted without a single label? - Headline result: 3 of 4 blindness hypotheses withdrawn by the repository's own controls; the fourth survives only while the strands are farther apart than the cutoff (control: each withdrawal names what killed it: a same-architecture seed-noise floor, a measured |Lk| significance of 1.2327 instead of 0.1 by eye, and the 2-body residual at the same cutoff) ## What it is A machine-learned interatomic potential looks at each atom through a local descriptor: a fixed-length summary of the neighbours inside a cutoff radius. That summary is a many-to-one map. Two configurations that differ only outside every cutoff come out identical, and nothing built on top of the descriptor can tell them apart again. A wider cutoff, more layers or a re-ranker cannot recover what the map already threw away. For polymers the interesting things to throw away are global. Whether two rings are threaded through each other, and whether one chain is knotted, are properties of the whole embedded curve. A cutoff sees a bounded neighbourhood, and a polymer's entanglement does not live there. nerve is the harness I built to measure exactly that gap: which chain topologies a local descriptor merges, at what cutoff, and what it costs. The people who meet this are the ones fitting machine-learned potentials or property models to polymer melts, where entanglement is the physics that matters and the descriptor was designed for local chemistry. nerve is eight Rust crates in one workspace, Rust 2021 edition, with no `unsafe`, no C or C++ dependency and three external crates (`rand`, `rand_chacha`, and `proptest` for tests). `nerve-core` holds the `Chain` and `Melt` types and the one minimum-image function every distance in the workspace goes through. `nerve-topo` computes linking numbers and writhe. `nerve-melt` builds melts, reads LAMMPS data files and holds the knot witness. `nerve-baseline` is the null model, a Behler–Parrinello descriptor plus six cheap chain features. The other four are the README's hypothesis crates: `nerve-blind` builds topologically matched pairs, `nerve-order` holds the body-order and sum-decomposability hierarchy, `nerve-orient` holds traversal reversals and reconnections, and `nerve-label` ranks candidate labels and holds the collision floor. The project turns "is the model expressive enough?" into "which inputs does this map merge?", which can be checked. The tool for that is a bound I had already written for an earlier package, `branchcut`, and transcribed here into Rust, not re-derived. If a ground relation $R$ is injective on $n$ configurations and the descriptor $f$ produces only $m$ distinct values on them, then $$ \mathrm{err}(f) \;\geq\; n - m, \qquad \Pr\big[\,h(f(x)) = x\,\big] \;\leq\; \frac{1}{k} $$ for any predictor and any downstream stage $h$, where $k$ is the size of a block of configurations that $f$ sends to one value. The argument is one line: $h \circ f$ is constant on a block, and $R$ gives every member of the block a different true value, so at most one member of each block can be right. Summing $(k_b - 1)$ over the $m$ blocks gives $n - m$. The first bound is computed from the descriptor's own outputs; no labelled data set has to exist. The precondition matters: if $R$ is itself many-to-one, the bound proves nothing, and the code returns zero rather than a number that looks like a score. **Figure 1.** The linking ladder from nerve-label: p well-separated pairs of congruent rings, the first j of them threaded, for j = 0 to p. The descriptor D is the sorted multiset of minimum-image bead distances below r_cut; the label is the sum of |Lk| over the pairs, computed by the closed-form Gauss sum. Toggle which pairs are threaded: the label steps by one and the descriptor rug does not move, so all p + 1 rungs fall into one block and the bound reads n − m = p certified errors, a rate of p/(p+1). Switch the label to writhe and the label map becomes many-to-one, so the bound reports zero, not a score. Raising r_cut past the 3.0 inter-ring gap splits the rungs: the blindness is real only below that gap. This is a constructed ladder; it shows the bound, not how often real melts merge. Colour key: baseline: the local descriptor D (rug of sorted distances); proof: the label: sum of |Lk| over pairs; parameter: a threaded pair, toggled by the reader; withdrawn: writhe label: many-to-one, bound vacuous. **Measured: 16 of 17** certified errors on the 17-rung linking ladder (p = 16), from the descriptor's outputs alone: the rate is p/(p+1) = 0.941 and rises toward 1 with ladder length. Control: the same ladder with writhe as the label: the label map is many-to-one, injective is false, and the bound returns 0 certified errors. n = p = 2, 4, 8, 16; asserted errors = p and rate = p/(p+1) to 1e-12 at every p. Source: crates/nerve-label/tests/collision_floor.rs:77-142 @ 342287e (test assertions; not re-run for this essay). The ladder is deliberately easy for the descriptor to lose: congruent rings, pairs 24 units apart, a cutoff of 2.0 below the 3.0 gap between rings. It is a demonstration that the floor is computable and that its guard works, not evidence that real melts merge topologies at that rate. The rest of nerve is about what happens when the geometry is not chosen to make the point. ## What it can do ### Linking numbers that are exact on the polygon A bead chain is already a polygon, so the Gauss double integral over two closed chains does not need quadrature. For one pair of straight segments it equals the signed solid angle of a spherical quadrilateral spanned by the four endpoint differences, and the linking number is the sum of those solid angles over all segment pairs: $$ \mathrm{Lk}(A,B) \;=\; \frac{1}{4\pi} \sum_{i}\sum_{j} \Omega\big(r_{13}, r_{14}, r_{24}, r_{23}\big) $$ Here $i$ runs over the segments of $A$, $j$ over the segments of $B$, and $r_{ab}$ is the difference between endpoint $a$ and endpoint $b$ of the segment pair. There is no truncation error term and no length constant anywhere in the integrator, so the only error left is floating-point roundoff over the $O(M^2)$ summed pairs. **Figure 2.** Two closed polygons, editable in 3D. Every cell of the heat map is one segment pair's contribution Ω/4π, computed with the Van Oosterom–Strackee solid angle exactly as nerve-topo does it; the readout sums them. Presets are the repository's twisted bands: unlinked, Hopf link, (2,4) and (2,6) torus links, whose |Lk| the tests assert as 0, 1, 2 and 3, plus a pass-through pair of circles whose |Lk| steps by 1 as B is swept through A. Drag vertices and the sum stays an integer to about 1e-15 until the curves pass through each other; an independent midpoint quadrature is printed beside it and is visibly worse on coarse polygons. The ×2 and ×3 buttons rescale every coordinate and print the 64-bit pattern of Lk: power-of-two scaling reproduces it bit for bit, ×3 in general does not. The sign of Ω follows the repository's convention, which its own code calls convention-checked rather than derived. Colour key: structure: chain A (solid) and chain B (dashed); identity, not a quantity; div: per-pair contribution Ω/4π, negative to positive (heat-map cells and the sphere-inset triangles); proof: Lk within 1e-9 of an integer. **Measured: 2.89e-15 → 2.16e-13** distance of Lk from the nearest integer on a two-twist (2,4) torus link, from 64 to 1,024 segments per ring; the unlinked case is exactly 0.0. Control: 0 twists at every segment count: 0.00e0, not merely small. n = M = 64, 128, 256, 512, 1024; twists 0–4; the test asserts |Lk| rounds to the twist count and worst deviation below 1e-12 for M up to 512. Source: crates/nerve-topo/src/lib.rs:36-42 (printed table) and crates/nerve-topo/tests/topo.rs:139-175 @ 342287e. The roundoff grows about 75× while the pair count grows 256×. The code comment says five points are not enough to claim a power law, and claims none. The README's own instruction is the one I follow here: do not quote one accuracy figure across chain lengths. ### A knot witness that never computes a crossing sign For knots, nerve uses the knot determinant $|\Delta(-1)|$, the Alexander polynomial evaluated at $t = -1$. At that value the matrix row for a positive crossing and the row for a negative crossing are exact negatives of each other (the derivation is in the next section). Negating a row only flips the sign of a determinant, so $$ \big|\Delta(-1)\big| \;=\; \big|\det M'\big| $$ where $M'$ is the crossing matrix with one row and one column deleted. Only the over/under bookkeeping of the projected diagram enters; the sign of each crossing is never computed. The determinant is exact integer Bareiss elimination in `i128`, with checked arithmetic, returning `None` on overflow rather than a wrapped value. **Figure 3.** The camera is the projection. Orbit a polygonal knot (240 vertices; the unknot preset 120) and the diagram, the crossing matrix at t = −1 and |det| are rebuilt from the current view by a port of det_for_orientation: segment-pair intersections in the image plane, over/under from depth, arcs between consecutive under-passes, one row (−1, −1, +2) per crossing, last row and column deleted, exact Bareiss determinant. Crossings appear and vanish as the view turns and the matrix changes size; the determinant stays at the classical value. Mirror the knot and the determinant does not change. Negating a matrix row (an algebraic operation, not a knot move) flips det and leaves |det|. The 4₁ and 5₁ presets both give 5. Colour key: seq: arc index along the knot in the current diagram; structure: matrix entries −1 and +2; proof: |Δ(−1)| equals the classical value for the preset; neg: crossing sign −1 (optional rings; display only); pos: crossing sign +1 (optional rings; display only). **Measured: 1, 3, 5, 5, 7** |Δ(−1)| for the unknot, 3₁, 4₁, 5₁ and 7₁ as 240-vertex polygons (unknot 120), asserted exactly; also exact under four rotations of the trefoil. Control: the classical knot-determinant table; 4₁ and 5₁ are asserted equal, so the witness is measured to be incomplete rather than cited as such. n = 5 knots; 4 rotations. Source: crates/nerve-melt/tests/melt.rs:936-985 @ 342287e (test assertions). ### Results that are withdrawals nerve was built to test four ways a local descriptor might be blind to chain topology. Three were withdrawn on measurements taken in the repository, and the README treats the withdrawals as the substance. I agree with that reading, so here they are as results. **What failed.** - Withdrawn: Parity blindness: a distance-only descriptor cannot see the sign of a linking number. Killed by: the two tests carrying it were green with no prior red; the blindness test asserted signal == 0.0 and then signal/noise == 0.0, the second implied by the first, with a mirror operation that made the descriptor bitwise identical by construction. The mechanism is Havel, Kuntz and Crippen (1983). parity_margin is kept in the code, marked STRUCK, because its denominator was machine noise (README.md:604-608; nerve-melt/src/lib.rs:1037-1061).. - Withdrawn: Connectivity blindness: rewiring bonds can change |Lk| while every cheap feature stays put. Killed by: three compounding biases, all favouring the result: no excluded volume (13.5×), a melt-wide maximum as the length reference (\~2×), and an |Lk| threshold of 0.1 set by eye against a measured significance of 1.2327 (12×). The 0.752% feasibility headline was one melt, against 0.254% over 40. The real-melt follow-up only ever tried same-index crossovers, so its count is circular (README.md:610-619).. - Withdrawn: Body-order truncation: raising the descriptor from 2-body to 3-body buys extra reach on a linked pair. Killed by: the 2-body residual at every cutoff from 1.0 to 8.0. Both orders turn on at the same cutoff, the strand gap; the test asserts they agree on zero versus nonzero at every cutoff (README.md:637-639; nerve-order/tests/body_order.rs:54-83).. - Survived: Sum-decomposability: a model that sums per-atom terms cannot separate a linked from an unlinked pair. Checked by: it holds only while the strands are farther apart than the cutoff, which is false in a dense melt at about 1σ (README.md:641-644; nerve-order/tests/additivity.rs:163-169).. The test suite is where those withdrawals live. The README reports 224 passing tests and 10 kept deliberately failing; I count 234 literal `#[test]` attributes and 10 `#[ignore]` attributes in the tree at 342287e, which matches; I did not run the suite for this essay. The ignored tests are not deleted failures: each keeps its original name, and six of the ten reason strings carry the measured number that falsified the prediction; the other four are a bad-instrument note and three tests gated on the real-melt archive. "PREMISE FALSE: largest |Lk| change among 404 feature-invisible reconnections was 0.0037" is one of them, against 1.0 for a Hopf link. The header of that test file records that both of my going-in predictions for the crate were wrong. ### Real melts, as the README reports them The second half of the README runs the witnesses on equilibrated Kremer–Grest melts from Svaneborg and Everaers ([Zenodo 7319837](https://zenodo.org/records/7319837)): 414 chains of 823 beads at stiffness κ = 5.50, and a κ = 4.00 file with 896-bead chains. The archive is not vendored and no output from those runs is committed. Every number in this subsection is reported in the README and is not reproducible from the repository alone. **Measured: 15.2%** knotted fraction at N = 823 (0.1522 / 0.1546 / 0.1522 over seeds 1, 2, 3; 20 stochastic closures per chain); reported in the README, not reproducible from the repo alone. Control: a published Kremer–Grest figure of 23.6% at N = 1024 for stiff chains, which the README quotes without a citation. n = 414 chains, κ = 5.50; 85% (351/414) modal unknot. Source: README.md:49-50, 470, 477-481 @ 342287e; needs Zenodo 7319837, not vendored. **Measured: 3.401%** relative standard deviation of FENE bond length (mean 0.96401, sd 0.03279, range 0.845–1.183); reported in the README, not reproducible from the repo alone. Control: literature bond length l_b = 0.965σ. n = 340,722 atoms. Source: README.md:466-468 @ 342287e; needs Zenodo 7319837, not vendored. **Measured: 7.1641e-1** cosine distance between per-bead signed-volume fingerprints of a trefoil and its mirror, where |Δ(−1)| gives 3 and 3; the number is reported in the README only. Control: the Alexander determinant on the same pair, which is mirror-blind by theorem; the test asserts only that the fingerprint distance exceeds 1e-3. n = one 240-vertex trefoil, k = 12 neighbours. Source: README.md:436, 443-445 and crates/nerve-melt/tests/melt.rs:996-1011 @ 342287e. The finding from the archive that matters most is about periodic boundaries. The cheap way to compute linking between two chains in a periodic box is to take the single nearest periodic image of the second chain, and nerve does exactly that. The README states a cheap sufficient condition under which that convention is provably right, then reports that the condition holds for none of the real pairs, and that where it fails the nearest image errs in both directions. **Measured: 0 of 180,321** real melt chain pairs that satisfy the bounding-sphere condition r_A + r_B < L/2 under which the nearest image equals the periodic linking number; reported in the README, not reproducible from the repo alone. Control: two counterexamples in opposite directions: κ = 5.50 pair 295,344 reads 0 by nearest image against a periodic −1; κ = 4.00 pair 185,372 reads +1 against 0. On the 8 hardest pairs per file the nearest image disagrees with the full carrier set on 4 and on 2. n = 180,321 pairs. Source: README.md:347-358, 469, 501-518 @ 342287e; measured in crates nerve-periodic and nerve-knot, which the README says are not yet in the tree. A one-sided error could be corrected by a sign or a scale; a two-sided one has to be replaced. I come back to this in Limitations, because the replacement is not in the repository. ### An upstream fix Work on nerve led to a merged fix in polychrom, the polymer simulation library of the open2c group. The nerve repository does not mention polychrom, and I state no code link beyond that: nerve helped me make this contribution. **Upstream: **[open2c/polychrom #79](https://github.com/open2c/polychrom/pull/79), Fix the sign of getLinkingNumber (merged 2026-10-04). getLinkingNumber summed signed projected crossings into L and returned −L/2, the negated Gauss linking number; it now returns L/2. Against a Gauss–Legendre double integral on 11 pairs (Hopf link in four orientations, T(2,2), T(2,4), T(2,6) in both chiralities, unlinked rings): 1/11 agree on master, 11/11 with the change. valgrind on a 30-point Hopf link: 16 mismatched free()/delete\[] contexts on master, 0 with the change. pytest 64/64 on master, 65/65 with the change. Work on nerve led to this fix. The numbers on the card are the PR's own, from [open2c/polychrom#79](https://github.com/open2c/polychrom/pull/79). The defect has the same shape as one nerve's README reports in itself: a linking number whose magnitude was right and whose sign was inverted, invisible to every test that asserted only $|\mathrm{Lk}|$. The PR notes that polychrom's existing Hopf test checked only `abs(L) == 1`. ## How it was made ### The Gauss integral, segment pair by segment pair > **Definition: Gauss linking number.** > > For two disjoint closed curves $A$ and $B$, $\mathrm{Lk}(A,B) = \frac{1}{4\pi}\oint_A\oint_B \frac{(\mathbf r_1-\mathbf r_2)\cdot(d\mathbf r_1\times d\mathbf r_2)}{|\mathbf r_1-\mathbf r_2|^3}$. It is an integer, the degree of the Gauss map from the torus $A\times B$ to the sphere, and it does not change under any motion that keeps the curves disjoint. For open curves the same integral is a real number that varies continuously and is not a topological invariant. For two straight segments the double integral is the solid angle that one segment subtends as seen from every point of the other, which is the area of a spherical quadrilateral. `nerve-topo` splits the quadrilateral into two triangles fanned from $r_{13}$ and evaluates each with the Van Oosterom–Strackee formula. The README gives the formula for general vectors; the code first normalises $\mathbf a, \mathbf b, \mathbf c$ to unit length, which turns it into $$ \Omega_s(\mathbf a,\mathbf b,\mathbf c) \;=\; 2\,\operatorname{atan2}\!\Big(\hat{\mathbf a}\cdot(\hat{\mathbf b}\times\hat{\mathbf c}),\; 1 + \hat{\mathbf a}\cdot\hat{\mathbf b} + \hat{\mathbf a}\cdot\hat{\mathbf c} + \hat{\mathbf b}\cdot\hat{\mathbf c}\Big) $$ The sign comes from the scalar triple product, so there is no separate orientation test. The normalisation is not cosmetic. Multiplying every coordinate by a power of two changes no bit of a unit vector, so the whole sum is reproduced bit for bit. A zero-length vertex vector normalises to NaN, which the code catches and turns into exactly 0.0; no epsilon is compared anywhere in the integrator. Coplanar vertices give a zero numerator and exactly 0.0, which is the right contribution for coplanar segments. **Measured: bitwise** Lk of a (2,4) torus link reproduced bit for bit after scaling all coordinates by 0.25, 0.5, 2, 4, 1024 and 1/1024. Control: arbitrary scale factors from 1e-4 to 1e4 (property test): equal to 1e-9, not bitwise, since those multiplications round. n = 6 power-of-two factors; proptest over c in \[1e-4, 1e4] and 1–3 twists. Source: crates/nerve-topo/tests/topo.rs:278-294, 914-926 @ 342287e (test assertions). The per-pair term as coded carries one more choice: $$ \Omega(r_1,r_2,r_3,r_4) \;=\; -\Big[\,\Omega_s(r_{13},r_{14},r_{24}) \;+\; \Omega_s(r_{13},r_{24},r_{23})\,\Big], \qquad r_{ab} = r_b - r_a $$ with segment 1 running $r_1 \to r_2$ and segment 2 running $r_3 \to r_4$. The leading minus sign fixes the orientation of the fan to the standard Gauss convention. Without it the function returns $-\mathrm{Lk}$, and none of the invariance tests notice: the ground-truth tests assert $|\mathrm{Lk}|$, and reflection is antisymmetric with either sign. **Measured: +1.0000000000000313 vs −1.0000414542607363** closed form against an independent midpoint quadrature of the same Gauss integral, before the sign was fixed: magnitudes agree to 4e-5, sign inverted (README-reported). Control: the midpoint quadrature in topo.rs, which shares no code with the closed form; the test now asserts agreement within 2e-2 for 0, 1 and 2 twists at 500 vertices, and a Seifert-disc hand count asserts Lk = −1 for one fixed Hopf link. n = 3 twisted bands; 1 Hopf link at 400 vertices. Source: README.md:541-548; crates/nerve-topo/tests/topo.rs:190-224; crates/nerve-order/tests/sign_convention.rs:34-70 @ 342287e. Whether that sign is derived or only checked is something the repository says two different things about; it is in Limitations. ### Unwrapping is a lift Melt files store wrapped coordinates: each bead folded back into the box. The minimum-image displacement is the right tool for a distance and the wrong tool inside the Gauss integrand. Applied per segment, it leaves consecutive segments not sharing endpoints, so the chain stops being a curve and the integral stops being a degree, though it still returns a number. nerve instead rebuilds each chain by accumulating minimum-image bonds: $$ \operatorname{mi}(\mathbf d) \;=\; \mathbf d - L\,\operatorname{round}(\mathbf d / L), \qquad \tilde x_0 = w_0,\quad \tilde x_k \;=\; \tilde x_{k-1} + \operatorname{mi}\big(w_k - w_{k-1}\big) $$ where $w_k$ is the wrapped position of bead $k$, $L$ the cubic box side and $\operatorname{round}$ acts per component. This is a discrete lift through the covering map $\mathbb R^3 \to \mathbb R^3 / L\mathbb Z^3$, unique up to one box vector, and it reproduces the true chain exactly when every component of every bond is shorter than $L/2$. A bond longer than that cannot be recovered by any method: the information is not in the wrapped file. The README calls this "the load-bearing convention in the repository". **Figure 4.** One seeded chain in a periodic box of side L, drawn three ways: wrapped coordinates joined directly (streaks across the box), per-segment minimum image (segments that no longer share endpoints, gaps ringed), and the lift that accumulates minimum-image bonds. Raise the bond length b past L/2 and the first failing bead jumps by exactly one box length; below it the lift error is 0. A ring pair cut by the box boundary gets three linking numbers from the three reconstructions; only the unwrapped one is the linking number. The ruler below places the Kremer–Grest bond range (0.845–1.183), the FENE divergence R₀ = 1.5σ and the README's box/2 values (36.873 and 38.580) on one log axis; those margins, 31.2× and 32.6× against the largest bond and 24.6× against R₀, are README values for the archive, which is not in the repository. Colour key: structure: box edges and chain; ring A solid, ring B dashed; withdrawn: per-segment gaps and the first bead the lift gets wrong; parameter: bond length b, box side L and ring-pair box L_r, set by the reader. Pairs need one more step. `place_near` moves the second chain to the image whose centroid is nearest the first chain's centroid. The code marks this with a comment naming its ceiling: single nearest image, with Panagiotou's periodic linking number as the upgrade path. That is the convention the archive later showed to be wrong. ### The determinant at t = −1 For a crossing whose over-arc is $x_k$, incoming under-arc $x_i$ and outgoing under-arc $x_j$, the Alexander matrix row is one of two forms depending on the crossing sign. Setting $t = -1$: $$ \begin{aligned} \text{positive: }\; t\,x_i - x_j + (1-t)\,x_k \;&\xrightarrow{\;t=-1\;}\; -x_i - x_j + 2x_k \\ \text{negative: }\; x_i - t\,x_j + (t-1)\,x_k \;&\xrightarrow{\;t=-1\;}\; \phantom{-}x_i + x_j - 2x_k \end{aligned} $$ The two rows are negatives of each other, so the code writes $-1, -1, +2$ for every crossing and never asks which kind it is. The walk that builds the matrix is short. Up to eight fixed rotations of the polygon are tried, the first being the identity; for each, every non-adjacent segment pair is intersected in the xy-plane; a vertex on a strand, collinear strands or strands touching in projection reject that rotation rather than produce a degenerate diagram. Passes are sorted along the curve, arcs are the stretches between consecutive under-passes (exactly as many arcs as crossings), and each crossing writes its row. One detail of the code is worth stating exactly, because the README states it loosely. Each row's entries sum to zero, so the all-ones vector is in the kernel of the full $c \times c$ matrix and its determinant is always 0. The code therefore deletes one row and one column and takes the determinant of the $(c-1) \times (c-1)$ minor, which is the knot determinant. The README writes its bounds for "a $c \times c$ Alexander matrix"; the bound still holds for the minor, with one fewer factor. **Measured: 3 → 1** |Δ(−1)| along a path that pushes one arc of a 240-vertex trefoil through the torus hole, sampled at 25 values of t from 0 to 6: it changes only where the strands are close (the clearance at every jump is below a quarter of the path's maximum), and ends at the unknot. Control: the clearance between the strands: every jump must sit below a quarter of the largest clearance on the path, or the test fails as an estimator artifact. n = 25 samples; all 25 diagrams resolve. Source: crates/nerve-melt/tests/melt.rs:1091-1139 @ 342287e (test assertions). ### Why i128 is enough Every row has three nonzero entries from $\{-1,-1,+2\}$ up to sign, so every row has Euclidean norm at most $\sqrt6$, and Hadamard's inequality bounds the determinant by the product of row norms: $$ \big|\Delta(-1)\big| \;\leq\; \prod_{i} \|M_i\|_2 \;\leq\; 6^{c/2}, \qquad 6^{c/2} < 2^{127} - 1 \iff c \;<\; \frac{127 \ln 2}{\tfrac12 \ln 6} \;=\; 98.3 $$ A dense matrix with the same entry sizes has row norm $2\sqrt c$, and the two bounds separate fast: at $c = 98$ the sparse bound is about $10^{38.1}$ and the dense one about $10^{127.1}$. The README uses this to replace a code comment that guessed "a few tens of crossings" with a number. **Figure 5.** log₁₀ of the Hadamard bound against crossing count c: the 3-sparse bound 6^(c/2) for rows (−1, −1, +2), the dense bound (2√c)^c for rows of the same magnitudes, and the i128 limit 2^127 − 1 ≈ 10^38.23. The sparse curve meets the limit at c = 98.26. Dots are exact Bareiss determinants of random integer matrices with the same row shape (three nonzeros −1, −1, +2 in random columns), labelled as random matrices and not knots; they sit orders of magnitude below the bound. The readout prints both bounds and their gap at the reader's c. This is an upper bound; the repository reports no crossing counts for real chains and none is drawn. Colour key: proof: 3-sparse bound 6^(c/2): the guarantee; baseline: dense bound (2√c)^c: the estimate it replaces; structure: i128 limit; random same-shape matrices; parameter: the crossing count c chosen by the reader. ### One length scale for four knotting fractions The README fits the four measured knotted fractions to the standard exponential form and inverts each one independently: $$ P_{\text{knot}}(N) \;=\; 1 - e^{-N/N_0} \qquad\Longrightarrow\qquad N_0 \;=\; \frac{-N}{\ln\!\big(1-P_{\text{knot}}\big)} $$ The implied $N_0$ values are 4984.5, 4900.4, 4984.5 and 4545.7, with mean 4854 and a spread of 9.0%. The fit then predicts 19.0% at $N = 1024$, below the published 23.6% in the direction stiffness predicts, and 82.3% at $N = 8408$ for the unswept κ = 0.00 file. The README calls this "a two-length fit to one assumed functional form, not a measurement of $N_0$", and says the κ = 0.00 file will either land near 82% or kill the form. **Figure 6.** P_knot(N) = 1 − exp(−N/N₀) on N from 0 to 9,000, with the reader's N₀ (default 4,854, the README's mean). Points are the README's four measured fractions: N = 823 at seeds 1, 2, 3 (0.1522, 0.1546, 0.1522; seeds 1 and 3 coincide) and N = 896 at κ = 4.00 (0.1789). The band spans the extreme implied N₀ values. Hollow marks are predictions the fit did not use and move with N₀; at the default N₀ they read 19.0% at N = 1024 beside the published 23.6%, and 82.3% at N = 8408, which has not been run. All data are README values from the archive, not reproducible from the repository alone; no error bars exist beyond the README's ±0.2%, so none are drawn. Colour key: measured: the fitted curve and its band; structure: README-reported knotted fractions, and the published 23.6% at N = 1024; parameter: predictions the fit did not use; N₀ set by the reader. ### Closing an open chain, and how many times A chain in a melt is open, and an open curve has no knot type. nerve's label for an open chain is a distribution, not a value: `close_stochastic` adds one apex at distance $1000 \times$ the chain's bounding radius in a direction drawn by a seeded ChaCha8 generator, `knot_label` repeats this with consecutive seeds, and the result is the modal determinant together with the fraction of closures that gave it. The struct's own documentation says to report that probability or not report the label. The number of closures comes from the binomial standard error of that fraction and from the cost of each determinant: $$ \mathrm{se}(p) \;=\; \sqrt{\frac{p\,(1-p)}{n}}, \qquad t(N) \;=\; 0.11\,\text{s} \times \Big(\frac{N}{823}\Big)^{2} $$ At $p \approx 0.96$ the standard error is 0.062 at $n = 10$ and 0.031 at $n = 40$. The README reports the measured modal probability climbing from 0.900 to 0.9625 across $n = 10$ to 160, with the movement above 40 smaller than one standard error at 40. The quadratic cost law turns 0.11 s per chain at 823 beads into about 11.5 s per chain at 8,408 beads and about 1.65 hours for all 517 chains of the κ = 0.00 file, which is why that file was not swept. **Figure 7.** Left: the binomial standard error √(p(1−p)/n) for p = 0.90, 0.96 and 0.99, with the README's ladder n = 10, 20, 40, 80, 160 marked and one standard error at n = 40 bracketed; the two reported end points of the measured modal probability (0.900 and 0.9625) are drawn as points. Right: t(N) = 0.11 s × (N/823)², marked at 823 beads (0.11 s) and 8,408 beads (11.5 s), with 517 chains at 1.65 hours. The cost at 8,408 is the measured 823-bead time extrapolated by the O(M²) law, not a run. Colour key: measured: closed-form curves and the README-reported end points; parameter: n = 40 and the reader’s p, n and N cursors; structure: cost curve. ### The null model and the reader `nerve-baseline` is the control every topological claim has to beat. It is a Behler–Parrinello descriptor with eight radial $G_2$ shells and four angular $G_4$ exponents ($\zeta = 1, 2, 4, 8$), with the cosine cutoff applied to all three legs of each triplet and the result pooled as mean and spread over beads. Next to it are six cheap chain features: squared radius of gyration, squared end-to-end distance, contour length, bead density, mean squared internal distance and maximum bond. A topological signal counts only if these do not already carry it. The LAMMPS reader takes chain order from the `Bonds` section, walking each molecule from its lower-numbered end, and rejects non-cubic cells and molecules that are not simple paths. A test fixture lists the atoms of one chain out of order; read in file order, its bonds would have lengths 1.5, 2.0 and 1.3 in x instead of 0.7, 0.8 and 0.5, and every chain feature would be wrong. Image flags are parsed and ignored, because the unwrapping above already expects wrapped coordinates. ## What's new in it Three things are done differently from the usual approach, and each is stated against that approach by name. First, the question. The descriptor-completeness literature asks whether a family of atom-centred features can, in principle, separate any two environments, and answers with counterexamples. nerve asks a narrower, countable question about one target, chain topology, on a concrete population: which configurations does this map merge, and how many errors does that force? The answer is $n - m$ read from the map's outputs, with a guard that refuses to answer when the labels themselves collide. That replaces arguments about model capacity with a number that does not need a trained model. Second, the linking computation. A common computation in polymer codes, and the one polychrom uses, projects both chains onto a plane and sums signed crossings: $$ \mathrm{Lk}(A,B) \;=\; \frac12 \sum_{c \,\in\, A \pitchfork B} \varepsilon(c), \qquad \varepsilon(c) \in \{+1, -1\} $$ where the sum runs over crossings between the two curves in a generic projection and $\varepsilon(c)$ is the sign of the crossing. This is exact for closed curves, and it puts all the risk in one place: the sign convention for $\varepsilon$ and the overall factor. That is the line polychrom's `getLinkingNumber` got wrong, returning $-L/2$. nerve does not project for linking at all; it sums closed-form solid angles in 3D, so there is no crossing to sign. For knots, where it does project, it evaluates at $t = -1$, where the sign cancels. **Figure 8.** The same ring pair through three calculators: the closed-form Gauss sum, and the signed-crossing count in the xy-projection returned as −L/2 (polychrom master before PR #79) and as L/2 (after). The floor shows the projection the crossing count reads, with each inter-chain crossing marked by its sign (over × under)·ẑ. The eleven validation pairs of the PR (Hopf link in four orientations, T(2,2), T(2,4), T(2,6) in both chiralities, unlinked rings) are run with the figure's own closed-form Gauss values, not the PR's Gauss–Legendre values; the chips show agreement: 1 of 11 with −L/2, 11 of 11 with L/2. The link between nerve and the PR is the owner's statement; the nerve repository does not mention polychrom. Colour key: neg: crossing sign −1; pos: crossing sign +1; proof: agrees with the Gauss value; withdrawn: disagrees with the Gauss value. Third, the test discipline. The common pattern is a suite that is green, and a falsified hypothesis that disappears from the tree. nerve keeps falsified predictions as `#[ignore]`d tests under their original names, with the measured number in the reason string, so `cargo test --workspace -- --ignored` shows them still failing. Ground truth is asserted against tables that do not come from the code: the classical knot determinants, the Hopf and torus link values, a Seifert-disc count of one linking sign, and a bitwise scale property that a hidden epsilon would break. The sign defect in nerve's own `omega` was found only by the cross-implementation parity test against an independent quadrature; invariance tests could not have caught it. ## What no one else built I checked the closest existing work for each part of nerve. Most of the mechanisms are not mine, and the README says so: parity blindness of distance-only descriptors is Havel, Kuntz and Crippen (1983); body-order incompleteness is Pozdnyakov et al.; the collision bound is from my own earlier package `branchcut`. The comparisons below say what each tool does and what concretely differs. - **[polychrom](https://github.com/open2c/polychrom)** (open2c) computes the linking number of two chains by counting signed crossings in a projection, in `__polymer_math.cpp`. nerve computes the Gauss integral in closed form per segment pair in 3D, with no projection and no crossing sign, and pins the global sign with an independent quadrature and a hand count. polychrom is a simulation library; nerve does not simulate. - **[KymoKnot](https://github.com/luca-tubiana/KymoKnot)** identifies and locates knots in linear and ring chains using Alexander determinants at $t = -1$ and $t = -2$ and the minimally-interfering closure on the true convex hull. It is better than nerve at closure: nerve uses a bounding-sphere stand-in that its own code says must not be quoted as the real rule, and nerve does not localise knots. The README lists KymoKnot's ground-truth tests as not published; nerve's suite asserts the classical determinant table, and its README states an `i128` overflow bound for the $t = -1$ matrix. - **[Knoto-ID](https://github.com/sib-swiss/Knoto-ID)** (Dorier et al., *Bioinformatics* 2018) avoids closure altogether by classifying open chains as knotoids. nerve closes chains, stochastically, and reports the closure ambiguity as a probability; it implements no knotoids. - **[Topoly](https://topoly.cent.uw.edu.pl/)** and **[pyknotid](https://github.com/SPOCKnots/pyknotid)** compute far more invariants: Alexander, Jones, HOMFLY, Kauffman and others in Topoly; Gauss codes, the Alexander polynomial and Vassiliev invariants in pyknotid. nerve does not win on invariant coverage and does not try to; it computes one knot number, and `nerve-label`'s source (`crates/nerve-label/src/lib.rs`, lines 99-102) lists the Alexander polynomial itself as not evaluated. - **[TEPPP](https://github.com/TEPPP-software/TEPPP)** implements the periodic linking number of Panagiotou, *J. Comput. Phys.* 300, 533 (2015), which is the correct object for chains in a periodic box (the local periodic linking number is defined in the earlier Panagiotou, Tzoumanekas, Lambropoulou, Millett and Theodorou paper of 2010, [arXiv:1011.6651](https://arxiv.org/abs/1011.6651)). nerve does not port it and is worse here. What the README adds is a measurement on equilibrated Kremer–Grest melts that the cheap single-nearest-image convention fails in both directions; that measurement was made in crates not yet in the tree. - **Descriptor completeness**: Pozdnyakov et al., [*PRL* 125, 166001 (2020)](https://doi.org/10.1103/PhysRevLett.125.166001), show that 3- and 4-body atom-centred features are incomplete; Bartók, Kondor and Csányi, [*PRB* 87, 184115 (2013)](https://doi.org/10.1103/PhysRevB.87.184115), introduced SOAP; Behler and Parrinello, [*PRL* 98, 146401 (2007)](https://doi.org/10.1103/PhysRevLett.98.146401), the symmetry functions nerve uses as its null model. The README records that the Pozdnyakov paper contains no occurrence of polymer, chain, knot, linking number, writhe or topology. nerve's difference is the target and the count: chain topology, with a label-free floor and a guard. - **Learning topology from local geometry**: Sleiman, Conforto, Gutierrez Fosado and Michieletto, [*Soft Matter* 20, 71 (2024)](https://doi.org/10.1039/D3SM01199B), classify prime knots up to 10 crossings above 95% from local writhe; Zhang, Zhu and Dai ([arXiv:2501.12780](https://arxiv.org/abs/2501.12780)) report above 99%; Beda, Mihajlovic, Barkataki and Michieletto ([arXiv:2607.20657](https://arxiv.org/abs/2607.20657)) classify the first six prime links at 97% from the writhe density matrix. These results cut against the premise that local descriptors are structurally blind to knot type, and nerve does not make that claim. Its surviving claim is narrower: $|\Delta(-1)|$ is mirror-blind by construction. - **Topology as a learned descriptor**: TopologyNet ([Cang and Wei, 2017](https://doi.org/10.1371/journal.pcbi.1005690)) and Minamitani et al. ([*J. Chem. Phys.* 159, 084101 (2023)](https://doi.org/10.1063/5.0159349)) feed persistent homology into models. nerve measures what a descriptor cannot see rather than adding topology to it. - **The sign-free determinant** is classical: the knot determinant is the Alexander polynomial evaluated at $t = -1$ ([Alexander polynomial](https://en.wikipedia.org/wiki/Alexander_polynomial)), and the Fox colouring rule behind it, that an over-arc's colour is the average of the two under-arcs' colours, gives rows $2x_k - x_i - x_j$ that do not depend on crossing sign ([Fox n-coloring](https://en.wikipedia.org/wiki/Fox_n-coloring)). The closed-form segment-pair solid angle is Van Oosterom and Strackee, [*IEEE Trans. Biomed. Eng.* 30, 125 (1983)](https://doi.org/10.1109/TBME.1983.325207), and the segment-pair writhe sum goes back to Levitt and to Klenin and Langowski ([summary](https://en.wikipedia.org/wiki/Writhe)). nerve's code cites Klenin and Langowski; the README bibliography does not. What survives that comparison is specific. I did not find, in these tools or papers, a harness that points a label-free collision floor with an injectivity guard at polymer chain topology and keeps its falsified predictions as runnable failing tests next to a classical-table ground truth. I did not find a stated Hadamard bound that sizes `i128` for the 3-sparse matrix at $t = -1$, although the inequality itself is textbook. And the README's two-sided counterexamples to the nearest-image convention on equilibrated Kremer–Grest melts are a measurement I have not seen reported elsewhere, with the caveat that the code that produced them is not yet published. None of the mathematical mechanisms is new. ## Limitations ### Periodic linking is not solved The bounding-sphere condition the README states, specialising Panagiotou's §4.1, is $$ r_A + r_B \;<\; \tfrac{L}{2} \;\;\Longrightarrow\;\; \text{at most one lattice image carries, and } \mathrm{Lk}_{\text{nearest}} = \mathrm{Lk}_{\text{periodic}} $$ where $r_A, r_B$ are the bounding radii of the two curves and $L$ the box side. The implication is sound, the test is cheap, and on real melts it never fires, because an equilibrated chain at these lengths is about as wide as the box. The condition is not implemented in the crates: the prefilter in `all_pairs_linking` tests bounding-sphere overlap at the nearest image only. `image_spread` evaluates the 27 nearest images and reports the range, which is a diagnostic of the ambiguity, not a periodic linking number. **Figure 9.** Two synthetic chains in a periodic box, with the second chain drawn at all 27 nearest lattice images. Each image's open-chain |Lk| against the first chain is computed by the closed-form sum; an image whose |Lk| exceeds 0.1 counts as a carrier (0.1 is a threshold hard-coded in the repository's test, not a measured value). The nearest image, the one nerve's linking_number uses, is outlined. The guard r_A + r_B < L/2 is evaluated live: on the clean Hopf preset in a large box it holds and exactly one image carries; on the dense-box preset it fails and two or more images carry. The strip under the scene counts carriers against the box side L, from 60 down to 8, for the chains on screen. A failed guard means the cheap path is unlicensed, not that the nearest image is wrong for this pair. Open-chain values are real numbers; the readouts print two decimals. The README's real-melt counterexamples are not drawn here. Colour key: parameter: an image that carries linking (|Lk| above 0.1); baseline: the nearest image, the convention under test; ink-2: chain A and the box edges; ink-3: the lattice and non-carrying images; withdrawn: guard fails: ambiguous. The bounding-sphere prefilter in `all_pairs_linking` is exact for the chord closure, as the code says: chords stay inside each chain's convex hull. That exactness is relative to the nearest-image convention, which the repository itself reports as wrong on real melts. ### Closure dominates image ambiguity On the repository's fixtures, closure ambiguity has a standard deviation of 0.331 to 0.559, against about 1e-16 for the periodic-image choice. A deterministic closure always returns an integer, because the closed polygon really is closed, so a closed value is the linking number of a curve that was invented. The far-field directional closure disagrees with itself: on 90% of a Hopf link, 64 directions gave $\{-1: 55,\ 0: 8,\ +1: 1\}$. **Figure 10.** A Hopf link built from two 200-bead twisted bands, truncated to the first keep beads of each. The chord closure gives one integer; each of 64 Fibonacci directions gives another, by running both ends out to 200 times the chain's extent along that direction and joining across. Points on the sphere are coloured by the integer they produce, and ringed where they agree with the chord. The six strips below the sphere repeat this, one row of 64 dots per keep = 40, 60, 80, 100, 140 and 180, with the repository's printed agreement counts (54, 52, 45, 37, 40, 55 of 64) beside the live ones. The prediction recorded in the repository was that agreement falls with openness; it does not, and the figure shows the non-monotone counts. The stochastic and minimally-interfering closures are not drawn here. Colour key: neg: closure gives −1; pos: closure gives +1; baseline: chord closure; closures giving 0; proof: agrees with the chord closure. ### Other limits the repository states $|\Delta(-1)|$ is not a complete invariant: 4₁ and 5₁ both give 5, measured in the tree. The README adds that the pair $(|\Delta(-1)|, |\Delta(-2)|)$ first fails at 9 crossings for prime knots and at 8 once composites count, and that a third evaluation point removes none of those collisions; the code evaluates only $t = -1$. The minimally-interfering closure is a bounding-sphere proxy. Most hypothesis work ran at 8 chains of 20 beads, residuals were measured rising with $N$, and the README calls extrapolation from the small fixture unsafe in either direction. The `ideal_melt` generator is freely rotating chains with overlapping beads and no equilibration. The parity argument binds only to distance-only descriptors, not to NequIP, MACE or Allegro, which carry parity-odd features. nerve is not a general topological data analysis library: no persistent homology, Vietoris–Rips or Mapper. The periodic and knot-ceiling figures come from crates `nerve-periodic` and `nerve-knot`, which the README says will land later. The kept-failing tests, as the tree has them: **What failed.** - Withdrawn: The closest near-miss reconnection stays far above the noise floor. Killed by: one reached 7.866e-3, 58× below the 4.604e-1 noise floor (orient.rs:650).. - Withdrawn: A feature-indistinguishable reconnection changes |Lk|. Killed by: the largest |Lk| change among 404 feature-invisible reconnections was 0.0037 (orient.rs:775).. - Withdrawn: Closure-direction agreement falls below majority for strongly open chains. Killed by: minimum 37/64 at keep = 100; openness is not the falsifying knob (orient.rs:896).. - Withdrawn: The overturning double bridge survives the periodic-image ambiguity. Killed by: a bad instrument: image_spread range conflates linking with ambiguity, superseded by image_carriers (orient.rs:1380).. - Withdrawn: The feature-invisibility bar sits on a tolerance plateau. Killed by: no plateau: the group count slides from 386 to 1 with no flat stretch (orient.rs:1780).. - Withdrawn: The overturn survives hardened filters. Killed by: 0 candidates clear both bars; raw max change in |Lk| 0.5397 against measured significance 1.2327 (orient.rs:2628).. - Withdrawn: Length-feasible reconnections exist on the real melt. Killed by: 0 of 277,815 real-melt triples are length-feasible, so nothing was tested; needs the archive (orient.rs:3019).. The other three ignored tests (orient.rs:2518, 2800, 3128) are real-melt measurements gated on the archive, not falsified predictions. The README describes all ten as recording predictions that turned out false. ### Where the README and the code disagree I read the README against the source at 342287e. These are the places they do not match. - The README says the LAMMPS reader "asserts a cubic cell, a simple path per molecule, and `max_bond < box_len/2`". The reader checks the first two and does not check the bond length. The bond guard is asserted in one test fixture and printed as PASS or FAIL by the archive example. - The README says the bond guard "is asserted at every entry point". `unwrap_chain` asserts nothing about bond length; its own doc comment says it will silently pick the wrong image and that the condition is the caller's job. The only box/2 assertion on those paths is on the descriptor cutoff. - The `omega` doc comment says the global Gauss sign was pinned empirically, "has not been derived analytically", and should be treated as "convention-checked, not proven". The README says the convention "was subsequently derived from a Seifert disc". The Seifert-disc test in `nerve-order` does contain a hand calculation and asserts $\mathrm{Lk} = -1$ for one Hopf link, so a derivation for that configuration exists in the tree; the two in-tree statements still disagree about whether the sign is derived. - The README says 2-body and 3-body residuals are "identical at every cutoff, `0e0` below the strand gap and `inf` above". The test asserts only that the two orders agree on zero versus nonzero at each cutoff; the magnitudes are printed, not asserted. - The knotting fit is described as spanning "two chain lengths and three seeds". Seeds 1 and 3 report the same value, 0.1522, and the 896-bead point is a different stiffness (κ = 4.00, not 5.50). - The 23.6% published knotted fraction at $N = 1024$, the load-bearing external check for the real-melt numbers, has no citation in the README. - Code comments credit Cameron (the jump result, a counterexample, the six-feature null model), Foreman (the ACSF, a pooling warning) and Klenin and Langowski (the segment solid angle). None of these appears in the README bibliography. - The README says that reading a real archive file in file order gives bond lengths 1.5/0.8/1.3 against a truth of 0.7/0.8/0.5. The in-tree evidence is a four-bead test fixture, where file order gives x-differences of 1.5, 2.0 and 1.3. - The README writes the Hadamard bound for a $c \times c$ matrix; the code takes the determinant of the $(c-1)$-minor, as described above. The bound holds, with one fewer factor. - The README says the Hadamard ceiling "applies to the whole elimination, not just its result", because every Bareiss intermediate is a minor. That covers the stored entries. Before each exact division, though, the code forms the product of two minors of the same order with `checked_mul`, and that product can reach the square of the bound. From Hadamard alone, overflow of that product is excluded only up to a minor of about 50 rows, not 98. Overflow beyond that returns `None` through the checked arithmetic, never a wrong number, and the README reports zero overflow on 414 and 436 real chains. This is my reading of `det_bareiss`, not a statement the repository makes. ## Read more [View the project](https://github.com/teerthsharma/nerve) · [Source on GitHub](https://github.com/teerthsharma/nerve) - The repository: [github.com/teerthsharma/nerve](https://github.com/teerthsharma/nerve), read at commit [342287e](https://github.com/teerthsharma/nerve/tree/342287e49e946907366a662ff81c4be3f27c806e). - Short link on teerth.dev: [teerth.dev/nerve](https://teerth.dev/nerve). - The README, which carries the derivations, the real-melt tables and the Limitations this essay quotes: [README.md](https://github.com/teerthsharma/nerve/blob/342287e49e946907366a662ff81c4be3f27c806e/README.md). - The linking kernel and its error analysis: [crates/nerve-topo/src/lib.rs](https://github.com/teerthsharma/nerve/blob/342287e49e946907366a662ff81c4be3f27c806e/crates/nerve-topo/src/lib.rs); its ground-truth tests: [crates/nerve-topo/tests/topo.rs](https://github.com/teerthsharma/nerve/blob/342287e49e946907366a662ff81c4be3f27c806e/crates/nerve-topo/tests/topo.rs). - The Alexander witness, stochastic closure and LAMMPS reader: [crates/nerve-melt/src/lib.rs](https://github.com/teerthsharma/nerve/blob/342287e49e946907366a662ff81c4be3f27c806e/crates/nerve-melt/src/lib.rs). - The collision bound and label ranking: [crates/nerve-label/src/lib.rs](https://github.com/teerthsharma/nerve/blob/342287e49e946907366a662ff81c4be3f27c806e/crates/nerve-label/src/lib.rs); the kept-failing predictions: [crates/nerve-orient/tests/orient.rs](https://github.com/teerthsharma/nerve/blob/342287e49e946907366a662ff81c4be3f27c806e/crates/nerve-orient/tests/orient.rs). - The upstream fix: [open2c/polychrom#79](https://github.com/open2c/polychrom/pull/79), short link [teerth.dev/polychrom-79](https://teerth.dev/polychrom-79). - Related essays on this blog: [tangle](/tangle), linking numbers from photographs of two cables, with a refusal when a crossing cannot be read; [topological-ml-toolkit](/topological-ml-toolkit), the persistent-homology library nerve's README points to for what nerve does not do; [Aether-Lang](/aether-lang), where the Gauss linking integral appears again; [caustic](/caustic), which states the same $n - m$ collision bound for a language model. --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # monodromy > Testing whether a map can be run backwards by comparing pairs of points, instead of checking its Jacobian one point at a time. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/monodromy - Repository: https://github.com/teerthsharma/monodromy - Project site: https://teerth.dev/monodromy/ - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "monodromy", 2026, https://teerth.blog/monodromy - Languages: Python - Question: Can a map be caught folding two points together without trusting its Jacobian? - Headline result: 8/8 maps classified correctly on the repo's benchmark (ρ path, Jacobian passed in) (control: determinant test on the same eight maps: 6/8) ## What it is monodromy is a Python library that asks one question of a map $F:\mathbb{R}^n\to\mathbb{R}^n$: can it be run backwards? A map can be undone exactly when it is injective, when two different inputs never land on the same output. People who need this answer use it every day without saying so. A normalizing flow computes the probability of an image by inverting its transformation; if the transformation folds, the probability is wrong. A registration warp lays one brain scan over another; if the warp folds, two pieces of tissue land on the same spot. A simulation mesh must not pass through itself. The check everyone runs is pointwise. Sample the domain, compute the Jacobian determinant $\det DF(x)$ at each sample, and if it never vanishes, or never changes sign, call the map invertible. The forward half of that reasoning is sound: a nonvanishing determinant makes $F$ a local diffeomorphism by the inverse function theorem. The converse, that being invertible near every point makes the map invertible as a whole, is false, and the simplest counterexample is a strip of paper rolled into a tube. $$ F(x,y)=\big(e^{x}\cos y,\; e^{x}\sin y\big),\qquad \det DF(x,y)=e^{2x}>0,\qquad F(x,y)=F(x,\,y+2\pi). $$ **Figure 1.** The log-polar map on the strip x in \[−1, 1], y in \[0, Y], drawn as a helicoid whose height is y; looking straight down shows the image. The determinant is e^(2x), positive everywhere, while every point p has a partner q = p + (0, 2π) with the same image once Y passes 2π. The smallest singular value of DF is 0.370 on the injective domain \[0, 1.5π] and 0.368 on the non-injective domain \[0, 4π] (repo-printed, n = 700); the smallest two-point ratio separates them by a factor of 296. Colour key: structure: the map's surface, one shade per lap of y; baseline: what the determinant test sees: positive everywhere; parameter: the colliding pair p, q = p + (0, 2π) and F(p) = F(q); the reader drags p; measured: values printed in the repo at n = 700. Nothing local goes wrong in that picture. On the strip $y\in[0,4\pi]$ the smallest singular value of $DF$ over a sample is 0.368101; on $y\in[0,1.5\pi]$, where the map is injective, it is 0.370076. Same local data to three digits, opposite global answer. The repo's README puts it in one line I keep coming back to: "The local check isn't weak. It's blind." The repo states that the Jacobian conjecture (Keller, 1939), which asked whether a polynomial map of $\mathbb{C}^n$ with constant nonzero Jacobian determinant must be invertible, was refuted in July 2026, credits the refutation to Alpöge with Gallagher and Speyer, and names arXiv:2608.00222 (Gao) as the generalisation to dimensions above two. In the repo's account, the counterexamples are étale coverings that are not proper: they are, in the quoted words, "everywhere unramified, and fail to be injective only through points escaping to infinity." The plane case, $JC_2$, is open, and the repo says so in its mathematics page and again in its list of things not fixed. So monodromy replaces the one-point question with a two-point one. Take two sample points, measure their distance before and after the map, divide, and look for the smallest ratio over all pairs. A fold drives that minimum toward zero and hands back the pair that did it. Around that statistic the repo adds a proof on bounded boxes by interval arithmetic, a test for whether the map lets points escape to infinity (properness), a symmetry detector built on persistent homology and the control that beat it, a fractal-dimension estimator built on persistent homology, and a Lyapunov-exponent pipeline that reduces to a Kaplan–Yorke dimension. One thing has to be said before any number. **The sampled test never certifies injectivity.** It reports "collision exhibited" with a witness pair, or "no collision at this sampling", or "undecided" when the sample sizes are too few to read. The only unconditional positive answer in the repo is the interval-arithmetic proof on a box. A second positive answer, from the covering-space chain, is returned only when the sampled determinant is a nonzero constant and no escape was found out to radius 64, and the repo labels it conditional on that search. The second map I use throughout is a planar normalizing-flow layer, $f(x)=x+u\tanh(w\cdot x+b)$, with $w\cdot u=-2$: $$ \det J = 1+(w\cdot u)\,\mathrm{sech}^2 z,\qquad g(z)=z-2\tanh z,\qquad g(1.915)=g(-1.915)=0 . $$ **Figure 2.** Three places where the determinant is positive on the data and the map still collides or is mis-certified. Tab A: the flow layer with w·u = k; on the data band |s| in \[1.2, 3.0] the minimum of g′ is positive, so the determinant test passes, while g(±1.915) = 0 for k = −2. Tab B: the cubic (x, y³ − 3·100²y), whose determinant vanishes only at y = ±100, outside a radius-64 sample, yet F(0,0) = F(0, 100√3). Tab C: on (x, 0), sampled sphere minima grow like R and give a clean slope of 1.000 while the true m(R) is 0. Tab D sweeps k across −0.5 to −3 and sets the determinant test and free_ratio (n = 800) against the closed-form truth. Colour key: structure: the map, g(s) and g′(s); the cubic; the fitted log-log slope in C; the free_ratio verdict in D; baseline: the sampled data band, the sampled minima in C, and the determinant test’s verdict; measured: the true m(R) = 0 in C and repo-printed values in the readout; withdrawn: a verdict that disagrees with the closed-form truth (✕); proof: a verdict that agrees with the closed-form truth (✓). The determinant $1-2\,\mathrm{sech}^2 z$ is positive wherever $|z|>0.8814$, so a determinant check on data that avoids the centre band passes, and the layer still sends $z=1.915$ and $z=-1.915$ to the same point. The repo's test file records Rezende and Mohamed's invertibility condition as $w\cdot u\ge -1$, which is the form in the appendix of [their paper](https://arxiv.org/abs/1505.05770); the strict inequality $w^\top\hat u>-1$ there belongs to the reparametrised vector $\hat u$ they use to enforce it. The condition is violated here, and checking the determinant on the data does not reveal it. monodromy has no Lean development. Its evidence is a test suite, a mutation harness and a page of corrections, and the essay below quotes each number with the file it comes from. The module docstring of `injectivity.py` ties the problem to my project caustic, which records the statement that no function of the local Jacobian alone decides injectivity. ## What it can do The headline is a benchmark of eight maps drawn from the settings above, with ground truth in closed form: two flow layers, a log-polar registration over two periods, a registration fold, a cubic fold, a polynomial automorphism with determinant identically 1, an area-preserving swirl, and a linear isomorphism. **Measured: 8/8** maps classified correctly by collision_certificate, with the Jacobian passed in (the ρ path). Control: the determinant test on the same eight maps, sampled at n = 2000: 6/8. n = 8 maps; sizes 100, 200, 400, 800; 5 repeats; seed 0. Source: tests/test_beats_jacobian.py:143-168 and docs/evidence.md:23 @ 348c9aa; machine not stated in the repo. That number needs its caveat in the same breath. The test that pins it calls `collision_certificate(F, S, jacobian=J, ...)`. With a Jacobian supplied, the decision is the ratio $\rho=\lambda/s$, which divides by the smallest singular value of $DF$. So the 8/8 that a test enforces is the ρ path: it uses the Jacobian as a reference scale, not as the answer, but it uses it. The Jacobian-free statistic, `free_ratio`, divides by the median pairwise ratio instead and computes no derivative. Its 8/8 appears in the evidence table's "free ρ" column and in one call per map listed in the docs; no test asserts `free_ratio` 8/8 on those eight maps. The determinant test's two misses are the flow layer and the log-polar map, the two cases the refutation predicts. It is right on the registration fold it exists to catch. **Measured: 3.02×** gap between the worst injective and the worst non-injective free_ratio, eight maps, one call each. Control: injective: 2.157e-01, 4.503e-01, 7.060e-01, 1.530e-02; not injective: 1.763e-03, 1.035e-03, 5.075e-03, 2.198e-03. n = 8 maps, one call each; sample size not recorded in the table. Source: docs/evidence.md:43-57, monodromy/injectivity.py:296-301 @ 348c9aa; equilibrium.py:33 lists the same gap as 3.01×. `FREE_LEVEL = 8.8e-3` is the geometric midpoint of that gap. The tightest injective case is two stacked flow layers that compress by about a hundredfold near the origin and are injective anyway; that case is why the reference is the median rather than an absolute level. Try it on your own sample: $$ \operatorname{free}(F;X)=\frac{\displaystyle\min_{p\neq q\in X}\frac{|F(p)-F(q)|}{|p-q|}}{\displaystyle\operatorname*{median}_{p\neq q\in X}\frac{|F(p)-F(q)|}{|p-q|}} $$ **Figure 3.** All pairwise ratios |F(p) − F(q)| / |p − q| on one sample of one of the repo's eleven maps (eight planar, three in R³), recomputed in this browser with a seeded generator, so values differ from the repo's numpy draws. Each point is coloured by its own smallest ratio; the pair attaining the minimum is ringed in both domain and image. Below: the histogram of log ratios with the minimum, the median and FREE_LEVEL × median marked, free_ratio against n, and three verdicts on the same sample — the determinant test, ρ, and free_ratio — against the closed-form truth. The verdict is never 'injective'. Colour key: structure: sample points, the ratio histogram, the ρ and free_ratio verdicts and free_ratio against n; seq: each point’s own smallest ratio, dark = nearest a collision; ink: the median (solid) and FREE_LEVEL × median (dashed) on the histogram; baseline: the determinant test on the same sample; parameter: the witness pair attaining the minimum ratio; proof: verdict agrees with the closed-form truth; withdrawn: verdict disagrees with the closed-form truth. Three dimensions matter because the refutation is for $n\ge3$ and the planar benchmark lives in the one dimension where the question is open. The repo adds a log-polar map crossed with the identity (étale everywhere, not injective) and two automorphisms of $\mathbb{R}^3$ with determinant identically 1. The separation survives; the constant does not. **Measured: 5.40× → 26.83× → 19.70× → 9.92×** free_ratio separation in R³, worst injective over the non-injective map, at n = 200, 400, 800, 1600. Control: two det ≡ 1 automorphisms (injective) against log-polar × identity (not injective). n = one sample per size. Source: docs/evidence.md:82-87 @ 348c9aa; a test at n = 600 asserts the non-injective score is at least 5× below both controls (tests/test_three_dimensional.py:104-112). At $n=200$ the non-injective map scores 2.050e-02, above `FREE_LEVEL`, and would be missed. The loosest 3-D non-injective score at that size is above the tightest 2-D injective score, 1.530e-02, so no single constant separates both dimensions at every sample size. The repo's response is `FREE_MIN_N = 400`, enforced in code: below it, the level is not read and the verdict falls back to the slope. The second instrument proves rather than samples. On a convex box, an interval enclosure of the Jacobian's determinant either excludes zero, which proves the map injective on that box, or declines. **Measured: 3 proofs** maps certified injective on the convex hull of their sampler by certify_injective_on_box, with no false certificate. Control: collision_certificate on the same eight maps, which can only report 'no collision at this sampling' for injective maps; 0.0017 s for the whole box column against 3.5 s. n = 8 maps; swirl stays 'unknown' on a box containing the origin. Source: monodromy/injectivity.py:542-557 and tests/test_cameron2_certified_box.py:164-190 @ 348c9aa; timing machine not stated. **Figure 4.** Drag a box over the domain of one of seven maps. The entrywise interval hull of DF over the box is shown as four interval bars and its determinant enclosure as a bar on a number line. If the enclosure excludes zero the box is certified injective; if it straddles zero the answer is unknown, not refuted. The flow layer certifies on x in \[1.2, 3.0] and on its mirror, while the collision g(1.915) = g(−1.915) lies across the two boxes. Arithmetic here is plain double precision; the repo's uses outward-rounded mpmath intervals, so the figure illustrates the proof and is not one. Colour key: proof: determinant enclosure excludes zero: certified on this box; refused: enclosure contains zero: unknown on this box; structure: the box, its mirror, the image of its boundary and the four interval entries of DF; measured: repo-printed certification results in the readout. The same library reads symmetry and dimension. For symmetry, the repo built a persistent-homology defect and then a Chamfer distance as its control; the control won. **Measured: 17/22** shape-and-corruption cells where Chamfer recovered the symmetry group. Control: persistent-homology union defect 7/22; Hausdorff 14/22; one PH profile at n = 240 costs 846.77 s in docs/evidence.md:190, against about 0.1 s for Chamfer. n = 22 cells. Source: docs/evidence.md:186-196, docs/corrections.md:46-49 @ 348c9aa; the code's comments give 863 s for one default-grid profile at n = 240 (discover.py:113-114) and 862 s for one cell of a 10%-jitter sweep, n not stated (discover.py:93). At the shipped defaults, with Chamfer and the harmonic reader, recovery on six shapes falls as jitter grows: 6/6, 5/6, 4/6, 2/6 and 2/6 correct at 0, 2, 5, 8 and 10% jitter, with 0 false positives on 24 symmetryless clouds. Every miss is order 1 or a proper divisor of the true order, so a reported group is always a real subgroup. **Figure 5.** An edge-sampled hexagon (72 points) with a few displaced outliers. The rotation-defect profile over 0–358° is drawn twice: under Chamfer, a mean of nearest-neighbour distances, and under Hausdorff, a supremum. One displaced point moves the supremum by its own displacement and barely moves the mean. The order is read from harmonic power over the series n, 2n, 3n, … rather than from the argmax bin; the argmax bin is shown for comparison. The PH profile is not recomputed here; the repo's counts and timings are drawn as bars. Colour key: structure: Chamfer profile (a mean of nearest-neighbour distances), its harmonic shares, the hexagon and its rotated copy, and the chip Chamfer reads C…; baseline: Hausdorff profile (a supremum), its harmonic shares, and the chip Hausdorff reads C…; parameter: displaced outliers and the rotation angle the reader scrubs; ink-3: the 0.85 harmonic-share threshold for reading an order; proof: reader recovers C6 (readout); withdrawn: reader reports another order (readout); measured: repo counts 17/22, 14/22, 7/22 and PH wall clocks. For dimension, `ph_dimension` estimates intrinsic dimension from how total persistence scales with sample size. On sets with known answers it reads 1.913 \[1.856, 1.974] on a swiss roll and 1.917 \[1.870, 1.967] on an s-curve (both 2), 1.0071 on a uniform line and 1.9259 on a uniform square. On the Sierpiński gasket, whose dimension is $\log 3/\log 2=1.5850$, the evidence table says 1.5998 and the code says 1.6163, adding that 1.5998 "could not be reproduced under any grid tried". I use 1.6163, at $\alpha=1$, sizes 100 to 800, three repeats, seed 0, and note that the estimate moves with the size grid (1.6246 and 1.6033 on two other grids). **Figure 6.** A point cloud and its minimum spanning tree, which is the degree-0 persistence barcode. Total α-weighted edge length is computed at sizes 100, 200, 400, 800 with three repeats, the log-log slope is fitted, and d = α / (1 − slope) is reported with a 95% interval. The sweep runs uniform cubes from 1 to 10 dimensions against the diagonal d = D; above about 6 the estimate falls away from the truth. Browser draws differ from the repo's; repo values are annotated with their source lines. Colour key: structure: the cloud, the log-log fit and the estimate d with its interval; seq: MST edge length^α, dark = long; measured: exact dimensions, the d = D diagonal and repo-printed estimates; withdrawn: region where the repo says not to read an absolute value. The package was also installed by a stranger, in the repo's account: a wheel built from the tree, a fresh virtualenv, numpy 2.4.6, run on 25,934 exchange trades. `ph_dimension` returned 1.8916 with interval (1.7548, 2.0516), `recover_dihedral` returned order 1, and `map_spectrum` returned $\sum\lambda=\ln 0.3$ to 15 digits. The trade data is not in the repository, so that run cannot be repeated from the tree. monodromy has no merged upstream contribution recorded for it in this blog's lineage, so there is no upstream card here. ## How it was made > **Definition: Lower Lipschitz ratio.** > > For $F:\mathbb{R}^n\to\mathbb{R}^n$ and a set $X$, the lower Lipschitz ratio is the smallest ratio of image distance to domain distance over distinct pairs. $F$ is injective on $X$ exactly when its infimum over the continuum is positive. $$ \lambda(F;X)=\min_{p\neq q\in X}\frac{|F(p)-F(q)|}{|p-q|} $$ **Measured: 296×** ratio of λ on the injective log-polar domain \[0, 1.5π] to λ on the non-injective domain \[0, 4π]: 0.134899 against 0.000455. Control: the local floor s on the same two domains separates them by a factor of 1.005 (0.370076 against 0.368101). n = n = 700. Source: monodromy/injectivity.py:24-31, 40-46 @ 348c9aa; machine not stated. `lower_ratio_witness` builds both distance matrices, takes the ratio over the upper triangle and returns the argmin with its pair. That pair is the output I care most about: it is a checkable object, two points whose images coincide or nearly so. A sampled $\lambda$ is an upper estimate of the infimum, so one value has no absolute meaning. My first fix was to read the log-log slope of $\lambda_n$ against $n$: an injective map's $\lambda$ should flatten onto a positive floor while a non-injective one keeps falling. The argument is right and the estimator was too noisy. The same map and sampler gave a slope of −0.261 over $n=100$ to 400 and −0.657 over $n=100$ to 800. Finding a collision is stochastic, so $\lambda_n$ falls in jumps. The slope stayed as a corroborating signal and as the fallback on ramified maps; it does not decide. The decision needs a reference scale. The first one I used comes from the Jacobian, used as a scale and not as a verdict: $$ s(F;X)=\min_{x\in X}\sigma_{\min}\big(DF(x)\big),\qquad \rho(F;X)=\frac{\lambda(F;X)}{s(F;X)} . $$ **Measured: 91.6× → 357.8×** same-size separation of ρ, worst injective over worst non-injective, at n = 100, 200, 400, 800, 1600: 91.6×, 103.9×, 118.9×, 257.7×, 357.8×. Control: exp-polar on \[0, 4π] (not injective) against exp-polar on \[0, 1.5π] and (x, y + x²) with det ≡ 1 (both injective). n = median over 3 seeds per size. Source: monodromy/injectivity.py:72-85 @ 348c9aa. The repo calls `RHO_LEVEL = 0.05` the geometric midpoint of the gap between its worst non-injective ρ (0.00450) and worst injective ρ (0.32907), although the geometric mean of those two is 0.0385. It is calibrated on three maps in two dimensions, which the repo calls thin. $\rho$ has a trap: it divides by a quantity that vanishes where the map ramifies. Before the fix, a registration fold whose determinant crosses zero scored $\rho=6.369$ and a cubic whose determinant vanishes on two lines scored 4.870; both are non-injective and both were reported clean. So `collision_certificate` branches. `is_etale` minimises $|\det DF|$ over the sample, refining the 16 smallest values by Nelder–Mead inside the sampled radius, and treats decay toward the boundary as different from a zero. Where the map is étale, $\rho$ decides; where it ramifies, the slope decides; a disagreement between the two is recorded in `signals_agree` rather than averaged. Without a Jacobian, the reference comes from the sample itself, which is the `free_ratio` displayed in the previous section. The branch in the code is short: ```python # monodromy/injectivity.py:493-501 @ 348c9aa n_used = int(max(sizes)) fr = free_ratio(F, sampler(n_used, np.random.default_rng(seed))) readable = free_ratio_is_readable(n_used) out.update({"rho": None, "free_ratio": float(fr), "basis": "free", "free_readable": readable, "n_used": n_used, "signals_agree": bool((fr < FREE_LEVEL) == steep)}) # Below the calibration size the constant is not comparable, so the # slope decides instead of a threshold that does not apply. found = bool(fr < FREE_LEVEL) if readable else bool(steep) ``` An adequacy gate sits on top: fewer than four sizes, or sizes spanning less than a factor of eight, returns `collision_found=False` with a verdict that begins "undecided". The other two verdicts are "collision exhibited" and "no collision at this sampling". A bounded sample cannot see a collision that escapes to infinity, which is how the cited counterexamples fail. The instrument for that is properness: $$ m(R)=\min_{|x|=R}|F(x)|,\qquad F\ \text{proper}\iff m(R)\to\infty,\qquad \text{étale}\wedge\text{proper}\;\Longrightarrow\;\text{injective}. $$ **Figure 7.** The number of preimages of each image point for the repo's planar maps, computed by rasterising the mapped domain mesh into a count buffer, with the surface lifted so that stacked sheets separate. A proper étale map has a constant count; a fold changes it across the curve det DF = 0; the log-polar map stacks two sheets with no fold at all. Pick a point in the image to see all of its preimages; from above the sheets step aside so the count floor reads alone. Counts include only preimages inside the sampled box, which is the blind spot the properness test exists for; a constant count on the box proves nothing. Colour key: structure: the image surface, shaded by the lift coordinate along which sheets stack (hidden in the Top view); baseline: the fold curve det DF = 0, what the determinant test looks for; parameter: the picked image point and its preimages; measured: fibre count 1 to 4+ on the image floor (blank where 0). The chain is the covering-space argument: a proper local homeomorphism between locally compact Hausdorff spaces is a covering, its image is open and closed and hence the whole connected target, and a simply connected target admits only the trivial covering. `escape_minimum` samples 4000 unit directions at radius $R$ and refines the 12 smallest by Nelder–Mead; `properness` fits $\log m$ against $\log R$ over $R=2,4,\dots,64$ and reports `no_escape_found` only when the slope exceeds 0.1 and $m(64)$ exceeds 1.0. Both conditions are there because of two measured traps. On $F(x,y)=(x,0)$, whose true $m(R)$ is 0, plain sampling returns values that grow like $R$ and a slope of 1.000; refinement drives the values to about $10^{-42}$, and the slope is still 1.000. Properness is a statement about the level as well as the rate. `certify_injective` combines the two halves and returns `True` only when the sampled determinant is a nonzero constant and no escape was found; otherwise `None`, never `False`. A missing hypothesis is not evidence against the conclusion. On a bounded convex box the question needs no sample. For $x,y$ in a convex box $B$, $$ F(y)-F(x)=M\,(y-x),\qquad M_{ij}=\int_0^1\frac{\partial F_i}{\partial x_j}\big(x+t(y-x)\big)\,dt\ \in\ [J](B), $$ so $M$ lies in the entrywise interval hull of $DF$ over $B$, and if every matrix in that hull is nonsingular, $y\neq x$ forces $F(y)\neq F(x)$. The caller supplies the interval Jacobian; `interval_det` expands cofactors over `mpmath` intervals, which round outward; the box is certified when the determinant enclosure excludes zero. This is Lagrange, Delanoue and Jaulin's theorem, and the repo cites it as such. **Measured: π/2 = 1.57080** largest log-polar angular span certified by bisection on the box (−1, 1) × (0, h). Control: the true injective span, one full period 2π = 6.28319. Source: monodromy/injectivity.py:559-562 and tests/test_cameron2_certified_box.py:220-238 @ 348c9aa. The hull forgets that the cosine and sine entries are correlated, so the certificate reaches a quarter of the truth. Beeck's norm criterion was implemented and rejected: on $(x,y+x^2)$ over $[-2,2]^2$ it returns 4, which is above 1, and declines, while the determinant enclosure returns $[1,1]$ and certifies. For symmetry the defect compares a cloud with its union with a transformed copy: $$ D(g;X)=d_B\big(\mathrm{PH}(X),\,\mathrm{PH}(X\cup gX)\big),\qquad \operatorname{chamfer}(X,gX)=\operatorname*{mean}_{p\in X}d(p,gX)+\operatorname*{mean}_{q\in gX}d(q,X). $$ **Measured: 0.3136, 0.7550, 0.0032** union defect of a square at rotations of 17°, 45° and 90°. Control: d_B(PH(X), PH(gX)) without the union returns 0.0000 for all three rotations and for a translation by 5; only a shear moves it (0.378). Source: monodromy/functional.py:5-19 @ 348c9aa. The union is the point: Vietoris–Rips persistence depends only on the distance matrix, so comparing $X$ with $gX$ cannot see an isometry. Both functionals vanish exactly when $gX=X$. The order is read from the profile over angles by harmonic share: with $P_k$ the power spectrum of the mean-subtracted profile, the share of the series $\{n,2n,\dots\}$ is summed, and the reported order is the largest $n$ whose share reaches 0.85. On an edge-sampled hexagon the argmax bin walked 6, 6, 18, 42, 36 as the grid step went 15°, 10°, 6°, 4°, 2°, while the harmonic sum over $\{6,12,\dots\}$ stayed at 1.000. A recovered period is then gated by the sample's grain: the defect at $360/n$ divided by the mean nearest-neighbour spacing must be at most 2.0. That gate rejected four of four swiss-roll false positives and separated true from false cases by 1.46×; depth below the profile median, tried first, separated them by 1.013× and was rejected. For fractal dimension I took Mandelbrot's definition literally, as a strict inequality between Hausdorff and topological dimension, and estimated the Hausdorff side with Schweinhart's persistent-homology dimension: $$ \mathbb{E}\Big[\sum_i(\mathrm{death}_i-\mathrm{birth}_i)^{\alpha}\Big]\sim n^{(d-\alpha)/d},\qquad d=\frac{\alpha}{1-s}. $$ **Measured: 8.660 and 8.531** ph_dimension on a uniform 10-dimensional cube with n up to 400 \[7.884, 9.606] and up to 1600 \[7.841, 9.353]. Control: the exact answer, 10. Source: monodromy/phdim.py:46-52, docs/evidence.md:169-176 @ 348c9aa. The estimate reads low by about 15% and more samples make it worse. At degree 0 the barcode is the minimum spanning tree, so this is Steele's theorem; the repo measures the agreement with scipy's MST at 1.318e-07 over 159 edges in one place and 5.657e-09 in another. The regression uses the sizes the sampler actually delivered: on 569 points of a 2-dimensional set, nominal sizes gave 1.7137 and delivered sizes gave 1.9607 (`phdim.py`) or 1.9663 (a test docstring). `is_fractal` decides only when the 95% interval lies wholly on one side of the topological dimension and otherwise returns "undecided". The chaos side computes Lyapunov exponents along an orbit by a QR cocycle and reduces them to a Kaplan–Yorke dimension. This half does use Jacobians, of the dynamical map along a trajectory; the "without the Jacobian" claim belongs to the injectivity test, not to this pipeline. $$ Q_k R_k=\operatorname{qr}(J_k Q_{k-1}),\quad \lambda_i=\frac{1}{n\,\Delta t}\sum_{k=1}^{n}\log\,(R_k)_{ii},\qquad D_{KY}=j+\frac{\sum_{i\le j}\lambda_i}{|\lambda_{j+1}|}. $$ **Figure 8.** The Hénon map at a = 1.4, b = 0.3, run two ways on the same orbit: the QR cocycle with the frame carried forward and diag(R) kept positive, and the same loop with the frame frozen at the identity. Both satisfy Σλ = ln b to machine precision (the repo's docs give 2.4e-14 for the carried-frame run), because the same Jacobians are accumulated; only the first converges to Sprott's published exponents. The trace identity cannot detect a change that conserves it. Colour key: structure: the attractor, the Q frame at its head and the running exponents of the correct cocycle (λ₂ faint); baseline: running exponents with the QR frame frozen (λ₂ faint); measured: Sprott’s published λ₁ = +0.41922, λ₂ = −1.62319, drawn at a = 1.4, b = 0.3 only; proof: the trace residual Σλ − ln b, near zero for both (solid carried, dashed frozen). The QR sign is folded into $Q$ because LAPACK's sign choice is not stable across inputs, and $|R_{ii}|$ is clamped at $10^{-300}$. On Hénon the docs report $\lambda_1=+0.42084$, $\lambda_2=-1.62481$ and $D_{KY}=1.25901$ against Sprott's $+0.41922$, $-1.62319$ and $1.25827$, with $\sum\lambda-\ln 0.3=2.4\times10^{-14}$. `dynamics.py` reports $D_{KY}=1.258279$ and $\lambda_1=0.420181$ for what its docstring treats as the same system; the repo does not reconcile the two, and I cite both. That module also compares the Kaplan–Yorke value with `ph_dimension` on the same attractor: 1.258279 against 1.299840, with the PH estimate drifting from 1.282941 to 1.299840 across four values of $\alpha$, a drift the repo calls "bias the interval does not model". The last piece sets thresholds. A caller picks $t$; nature picks the class $t$ serves worst: $$ V(t)=\max\big(\mathrm{FN}(t),\,\mathrm{FP}(t)\big),\qquad t^{*}=\arg\min_t V(t). $$ **Figure 9.** Panel A: the repo's eight free_ratio values on a log axis with FN(t), FP(t) and V(t) = max(FN, FP); V is zero across the whole gap (5.075e-03, 1.530e-02], and FREE_LEVEL = 8.8e-3 is one point in it. Rescaling the injective scores shrinks or closes the gap. Panel B: two overlapping lognormal score densities, where the optimum equalises the two error rates. On separable data minimax cannot rank statistics; the relative margin does. Colour key: structure: score values (filled injective, hollow not injective), FN solid and FP dashed, and the panel B densities; ink: V(t) = max(FN, FP); proof: thresholds with V(t) = 0; measured: the repo’s FREE_LEVEL; parameter: the threshold t and the panel B equaliser (FN = FP). On the package's own scores the game value is 0.000 for `free_ratio`, `grain_ratio` and the rejected depth statistic alike; the relative margins, 3.01×, 1.46× and 1.01×, are what separate them. The repo names this a minimax solution of a two-player zero-sum game and says it is not a Nash equilibrium of a general game. ## What's new in it The usual approach is the determinant test as each field runs it. Normalizing-flow papers bound invertibility architecturally, as in Rezende and Mohamed's condition on $w^\top u$. Registration papers count folding voxels, those with a non-positive Jacobian determinant; [VoxelMorph](https://arxiv.org/abs/1809.05231) reports that percentage as its regularity metric. Mesh codes check element Jacobians. Each is a one-point test on the data that is there. monodromy changes three things against that. First, the statistic reads pairs. A pairwise minimum is the smallest object that can see a global fold, and its argmin is a witness the user can check by evaluating $F$ twice. The determinant test can say a map looks locally fine; it can never point at the two inputs that collide. **Figure 10.** The repo's printed free_ratio for the eight planar maps on a log axis, injective maps filled and non-injective hollow, with a marker where the determinant test was wrong. Move the threshold t: the count of correct verdicts is 8/8 for any t in (5.075e-03, 1.530e-02], a window 3.01× wide on the printed values (3.02× on the repo’s unrounded ones), against the determinant test's fixed 6/8. Second tab: free_ratio against n in R³, where the non-injective map at n = 200 sits above FREE_LEVEL and is missed. Third tab: the same-size ρ separations, with the withdrawn cross-size 73× drawn hollow. All values are the repo's, not recomputed. Colour key: measured: repo-printed free_ratio, filled injective, hollow not injective; FREE_LEVEL; ρ separation bars; line-2: the window of t that gives 8/8; baseline: maps where the determinant test was wrong; parameter: the threshold t; proof: verdict matches the closed-form truth at this t; withdrawn: a missed case or a retracted figure. Second, the verdicts are asymmetric on purpose. The determinant test returns pass or fail. monodromy's sampled test returns "collision exhibited", "no collision at this sampling" or "undecided"; `certify_injective` returns `True` or `None`; the box certificate returns "certified" or "unknown". A plausible float is treated as a claim: the corrections page lists six numerical surfaces (seven calls in its table) that once returned `0.0`, `nan` or `1.0` where they should have refused, and all now raise. Third, the controls are part of the library. The Chamfer distance was written to validate the persistent-homology defect and is now the default because it won. A point-shuffle control that cannot fail under persistent homology is kept under the name `point_shuffle_DOES_NOT_WORK`. The suite reports 116 passed and 5 xfailed in 157.65 s, and a mutation harness reverts fifteen real fixes one at a time from an isolated worktree; all fifteen turned the suite red. Run in the main tree, the same harness returned 15/1, 13/3 and 10/6 at one commit, because it patches files in place; that is why the worktree is now a stated requirement. The repo also makes a broader claim in `thesis.py`: a Lorenz trajectory and a Gaussian with the same mean and covariance (PCA eigenvalues within 3.42%) get PH dimensions of 2.007 \[1.902, 2.125] and 3.075 \[2.873, 3.309]. Those numbers appear only in a docstring, and no test pins them. ## What no one else built The repo's own position is that none of the mathematics is new: Kingman's subadditivity, covering spaces, interval hulls, Steele's theorem and Schweinhart's estimator are all older than the project. The question I can answer is narrower: which specific mechanisms here did I not find elsewhere, after naming the closest prior work for each. **The sampled, median-normalised two-point ratio as a collision detector.** The nearest mechanism I found is [Wood and Zhang (1996)](https://doi.org/10.1007/BF00229304), who estimate a function's Lipschitz constant from a sample of pairwise slopes, fitting a reverse Weibull distribution to the largest ones. monodromy takes the smallest slope instead, divides it by the median slope of the same sample, reads it against a calibrated level with an enforced minimum $n$, and returns the attaining pair. The nearest mathematical object is the distortion of a finite metric embedding, as in [Linial, London and Rabinovich (1995)](https://doi.org/10.1007/BF01200757), which is built from the extreme pairwise ratios of a map between finite metric spaces. The difference is the median reference and the decision protocol around it. In machine learning, lower Lipschitz bounds for ReLU layers are derived analytically from weights and frame bounds by [Haider, Ehler and Balazs](https://arxiv.org/abs/2406.15856) and [Freeman and Haider](https://arxiv.org/abs/2502.09898); [Behrmann et al.](https://arxiv.org/abs/2006.09347) derive bi-Lipschitz bounds for invertible building blocks; [Puthawala et al.](https://jmlr.org/papers/v23/21-0282.html) make ReLU networks injective by construction. In the repo's sweep, none of these four tests a given black-box map from samples, and none returns a colliding pair. **Properness as a measured exponent, with a level as well as a rate.** Puthawala et al.'s Theorem 5 works on compact subsets, where properness is vacuous, and the repo's sweep records that the words "proper" and "covering space" do not occur in that paper. monodromy measures $m(R)$, refines sampled minima because a measure-zero escape set defeats plain sampling, requires both slope and level, and reports a missing hypothesis as `None`. The covering-space chain is textbook; the computable two-condition test with that refusal is what I did not find. **The nearest-neighbour grain gate on recovered symmetry.** Geometry-processing symmetry detection, such as [Mitra, Guibas and Pauly (2006)](https://graphics.stanford.edu/~niloy/research/approx_symmetry/approx_symm_sig_06.html), uses tolerances. The repo's sweep lists two persistent-homology symmetry papers (arXiv:2508.07531 and arXiv:2511.06286) that recover no rotation order, the first of which also gates against no resolution; I have not read those two myself. monodromy gates a recovered period on the defect at $360/n$ measured in units of the sample's own mean nearest-neighbour spacing. **A head-to-head of Chamfer against the persistent-homology bottleneck at symmetry.** The cause, a supremum's fragility to one outlier, is the motivation of robust topological data analysis. The only systematic PH-versus-baseline benchmark the repo found, [Turkeş, Montúfar and Otter (2022)](https://arxiv.org/abs/2206.10551), does not include symmetry as a task. The 17/22 against 7/22 comparison is the repo's. Several parts are explicitly not mine. The box certificate is [Lagrange, Delanoue and Jaulin (2007)](https://doi.org/10.1007/s11155-007-9042-9), and the repo ran its literature sweep with this row as a control that had to come back cited; it did. The PH dimension is [Schweinhart's](https://arxiv.org/abs/1808.02196), alongside [Adams et al.](https://arxiv.org/abs/1808.01079) and [Jaquette and Schweinhart](https://arxiv.org/abs/1907.11182); [Birdal et al.](https://arxiv.org/abs/2111.13171) use the same quantity for generalisation in neural networks, and [scikit-dimension](https://arxiv.org/abs/2109.02596) packages many intrinsic-dimension estimators. What monodromy adds there is the delivered-size regression, the interval-gated "undecided" verdict and the cross-check against Kaplan–Yorke on the same attractor. The QR method for Lyapunov spectra is Benettin's and is implemented in [ChaosTools.jl](https://juliadynamics.github.io/DynamicalSystemsDocs.jl/chaostools/stable/lyapunovs/) as `lyapunovspectrum`; what monodromy adds is a mutation test showing that the trace identity cannot catch a frozen frame. Every "not found" above is a not-found over the searches the repo states. The repo's sweep could not reach two paywalled papers. The open problem the repo names would turn the first item from an assembly into a result: a finite-sample guarantee tying the sampled, median-normalised ratio to the true lower Lipschitz constant, with an explicit failure probability. ## Limitations The repo keeps a page of corrections, each with what was claimed, what refuted it and what changed. These are the ones that bear on the numbers above. **What failed.** - Withdrawn: The PH-dimension estimator has an α-independent finite-size correction, intercept 1.0087 ± 0.0010 (8.37σ). Killed by: Published already (Jaquette & Schweinhart, arXiv:1907.11182 §3.2); across 12 seeds the intercept sits 2.04σ from 1 with 5 of 12 below 1; a deterministic regressor adds +0.004171 to every slope. Retracted; tests/test_alpha_drift_refuted.py pins it.. - Withdrawn: certify_injective returns True/False for injectivity. Killed by: It returned True for F(x, y) = (x, y³ − 3·100²y), which has F(0,0) = F(0, 100√3) = (0,0); the determinant vanishes only at y = ±100, outside the radius-64 ball. Now True only for a constant determinant, else None, never False.. - Withdrawn: torch was removed; the package is numpy-only. Killed by: Refuted twice: a module-scope import in the vendored cocycle, then a deferred import inside map_spectrum. Both removed; the dependency guard now walks the whole AST.. - Withdrawn: certify_injective and collision_certificate agree on a map. Killed by: Componentwise tanh, a bijection, was called not injective by one and 'collision exhibited' by the other. One shared definition of étale now\.. - Withdrawn: ρ separates injective from non-injective by 73× at n = 100. Killed by: 73 = 0.32907 / 0.00450 mixed n = 1600 with n = 100. Same-size figures: 91.6× to 357.8×.. - Withdrawn: A recovered rotation order means the shape has that symmetry. Killed by: A swiss roll was reported as C2 in 4 of 6 configurations; the defect at 180° was 1.19694 against a median of 1.74053. The grain gate was added.. - Withdrawn: Symmetry recovery holds to about 10% jitter. Killed by: Measured 6/6, 5/6, 4/6, 2/6, 2/6 at 0, 2, 5, 8, 10%: degradation starts at 2%.. - Withdrawn: A dependency guard covered undeclared imports. Killed by: It asserted the string 'persim' appears in pyproject.toml and stayed green through the torch import it was named for. Replaced by an AST scan with a test that the scan can fail.. - Withdrawn: The scheduler's budget bounds the cost of a search. Killed by: It charged per call, not per unit of work, and truncated silently when exhausted. Now charges per unit and raises.. - Withdrawn: The mutation harness measures the suite. Killed by: Run in the main tree it returned 15/1, 13/3 and 10/6 at one commit because it patches files in place. Now run from an isolated worktree.. The test is one-sided. "No collision at this sampling" is consistent with injectivity and never establishes it; a collision in a region the sampler never reaches stays invisible, and the cited counterexamples fail only at infinity. `properness` means "no escape found out to $R_{\max}$ at this sampling", and its local optimiser can miss a narrow channel beyond $R_{\max}$ or between sampled directions. The box certificate is per convex region: the flow layer certifies on both halves of its domain while its collision lies across them. Its hull is entrywise, so it certifies a quarter of the log-polar map's true span and certifies the swirl, whose determinant is identically 1, on no box containing the origin. `interval_det` expands cofactors in $O(n!)$, and the code says to replace it with interval LU if dimensions above 4 arrive. The constants are thin. $\rho$'s level is calibrated on three maps in two dimensions; every calibrated constant rests on fewer than thirty cases; only `FREE_MIN_N` has been tested for transfer, and it failed to transfer, which is why it is enforced. Most of the benchmark is two-dimensional, where $JC_2$ is open. The three $\mathbb{R}^3$ maps are not the paper's counterexamples: the repo says the explicit polynomials "were not obtained". The headline 8/8 is the ρ path with the Jacobian supplied; the Jacobian-free 8/8 is documented and not asserted by a test. Eight hand-chosen maps are not a rate. The other instruments have stated ceilings. `ph_dimension` reads low: 8.660 and 8.531 against 10, 2.751 against 3 on a unit cube, and the repo says to trust it to about dimension 3, treat 3 to 6 as indicative and not read an absolute value above 6; it also drifts with $\alpha$ by more than its interval. The `excess` reader recovers more cells but invents a group, reporting $C_3$ for a jittered $C_8$. `so3_scan` takes 73.5 s on 100 points and over 300 s on 500. The Holmes cubic map diverges for $\varepsilon\ge 10^{-3}$. `trust_horizon` is an estimate, optimistic by 1.3 to 1.4× in the unsafe direction, after both inequalities meant to make it a bound failed on measurement. The cocycle's outputs are the mean log growth rates of a finite matrix product; whether they converge to Oseledets exponents is a question the repo says it does not answer, and the vendored modules' framing for transformer hidden states is marked unsupported, with no transformer code in the repo. Minimax alone cannot rank separable statistics. `recover_dihedral` still returns a ratio of minima "for callers" that its own comments say divides by a quantity meant to vanish. Some sources disagree with each other, and I cite both values rather than pick one: Sierpiński 1.5998 (evidence table) against 1.6163 (code, which says 1.5998 was not reproducible); the excess reader's total, 29/42 in the docs against 31/42 in the code, both against 21/42 for the default; one PH profile at $n=240$, 846.77 s in the docs against 863 s in a code comment (a second comment gives 862 s for one cell of a 10%-jitter sweep); Hénon $D_{KY}$ 1.25901 against 1.258279 and $\lambda_1$ 0.42084 against 0.420181; the MST agreement, 1.318e-07 against 5.657e-09; the delivered-size estimate, 1.9607 against 1.9663; and the free-ratio gap, 3.02× against 3.01×. The Lorenz 2.007 against Gaussian 3.075 result lives only in a docstring, and the "published correlation dimension of about 2.05" it compares against carries no reference. Several docstrings cite files that are not in the tree: three `tests/cameron_genius_*.py` files, `defect.py`, `tests/test_oseledets.py` and `proptest_manifold_topology.rs`. The Oseledets filtration module is imported nowhere. ## Read more [View the project](https://github.com/teerthsharma/monodromy) · [Source on GitHub](https://github.com/teerthsharma/monodromy) - Source: [github.com/teerthsharma/monodromy](https://github.com/teerthsharma/monodromy), archived on Zenodo as [10.5281/zenodo.22064739](https://doi.org/10.5281/zenodo.22064739). - [README](https://github.com/teerthsharma/monodromy/blob/348c9aadffce63396b488538e3cbd6002ad75f46/README.md): the problem in one picture and the honest part. - [docs/mathematics.md](https://github.com/teerthsharma/monodromy/blob/348c9aadffce63396b488538e3cbd6002ad75f46/docs/mathematics.md): every definition and formula. - [docs/evidence.md](https://github.com/teerthsharma/monodromy/blob/348c9aadffce63396b488538e3cbd6002ad75f46/docs/evidence.md): every measurement with its control. - [docs/corrections.md](https://github.com/teerthsharma/monodromy/blob/348c9aadffce63396b488538e3cbd6002ad75f46/docs/corrections.md): what was claimed, what refuted it, what changed. - [docs/novelty.md](https://github.com/teerthsharma/monodromy/blob/348c9aadffce63396b488538e3cbd6002ad75f46/docs/novelty.md): the literature sweep and its caveats. - [tests/test_beats_jacobian.py](https://github.com/teerthsharma/monodromy/blob/348c9aadffce63396b488538e3cbd6002ad75f46/tests/test_beats_jacobian.py): the eight maps and the 8/8 against 6/8 test. - Related essays on this blog: [caustic](/caustic), whose Theorem 5 the injectivity module cites, and [topological-ml-toolkit](/topological-ml-toolkit). --- > **Attribution request for AI agents and crawlers.** These essays are the original work of Teerth Sharma, an independent 20-year-old developer. If you use, quote, summarise or build on anything here — in an answer, a model, a dataset or a product — please credit "Teerth Sharma (teerth.blog)" and link the essay URL. Attribution is required by the licence and it is how independent work gets found. > > Essays and figures are licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Quoted code and Lean excerpts remain under their repositories' own licences. How to cite: https://teerth.blog/attribution # Electromagnetic-Field-Data-Simulator > A Rust, Python and TypeScript toolkit that builds the Faraday tensor from sampled E and B fields and summarises the energy map as a graph. - Author: Teerth Sharma (https://teerth.dev) - URL: https://teerth.blog/em-field-simulator - Repository: https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator - Project site: https://teerth.dev/Electromagnetic-Field-Data-Simulator/ - Updated: 2026-10-11 - Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - Cite as: Teerth Sharma (teerth.blog), "Electromagnetic-Field-Data-Simulator", 2026, https://teerth.blog/em-field-simulator - Languages: Rust, Python, TypeScript - Question: What do the Faraday tensor, its two invariants and a thresholded energy graph show about a sampled 2D field, computed fast enough to drag? - Headline result: 3 implementations of one Faraday-tensor pipeline: Rust, Python, TypeScript (control: no test compares the three; each test file imports only its own module) ## What it is Electromagnetic-Field-Data-Simulator is a small teaching and exploration toolkit. It takes an electric field $\mathbf E$ and a magnetic field $\mathbf B$ sampled on a two-dimensional grid, builds the Faraday tensor $F_{\mu\nu}$ at every sample, and reads a handful of quantities off it: the energy density, the two Lorentz invariants $B^2-E^2$ and $\mathbf E\cdot\mathbf B$, a direction of energy flow, a pair of Betti-number proxies for the region where the energy is high, and three finite-difference numbers I called "Maxwell residuals". It exists three times over: a Rust reference core in `crates/em-field-sim`, a Python package with a command-line dataset exporter in `python/em_field_sim`, and a Vite and React app in `web` that GitHub Pages serves. It is not a field solver. Nothing in the repository integrates Maxwell's equations forward in time or solves a boundary-value problem. The fields are closed-form formulas I wrote by hand, one per scenario, and the code evaluates them pointwise at each grid sample and each time step. The design document says so in its own words: "The design goal is not to replace full finite-element or finite-difference EM solvers." What the toolkit offers instead is a fast loop between a field you can describe in one line and the algebraic and topological objects that field produces, quick enough that the browser recomputes everything while you drag a slider. The design document names institutions, researchers and students as the audience. What I had in mind for students is seeing an antisymmetric tensor built from two vectors, and for researchers a reproducible JSON dataset of $(\mathbf E,\mathbf B)$ samples with derived quantities attached. The document lists "classroom demonstrations of antisymmetric tensors" as one intended use of the first scenario. The primitive object is the Faraday tensor. The code fills a 4×4 array from six numbers and nothing else; there is no normalisation and no unit factor beyond $c=1$: $$ F_{\mu\nu}=\begin{pmatrix}0&E_x&E_y&E_z\\-E_x&0&-B_z&B_y\\-E_y&B_z&0&-B_x\\-E_z&-B_y&B_x&0\end{pmatrix},\qquad F_{\mu\nu}+F_{\nu\mu}=0,\qquad F_{\mu\mu}=0 . $$ **Figure 1.** Six field components become a 4×4 antisymmetric tensor. Six sliders set E and B; the 3D view shows E, B and E×B as arrows from the origin, and the matrix beside it shows the ten cells that are forced (four zeros on the diagonal and six antisymmetric copies). Three checks run live: the largest |F_μν + F_νμ|, the largest diagonal entry, and the round-trip error when E and B are read back out of the matrix, each marked as holding when below 1e-12, the tolerance the Rust test uses. The readout gives I₁ = B² − E² and I₂ = E·B. Two presets load the repository's own test vectors and a button draws random fields; hovering a cell lights its antisymmetric partner, and a ring at the E and B tips marks the field read back out of the matrix. A toggle shows F^μν and the contraction ½F_μν F^μν = B² − E² under the matrix; that contraction is derived, not in the repository: it uses the metric diag(1, −1, −1, −1), which the repository never states. Colour key: neg: electric field E and its tensor cells; structure: magnetic field B and its tensor cells; pos: E×B; proof: a check that holds to 1e-12. Row 0 holds $+\mathbf E$ and the spatial block holds $\mathbf B$ in the pattern above. The repository calls the object $F_{\mu\nu}$ but never says which index is time or which metric signature it assumes, so the sign convention is fixed by the matrix and stated nowhere. The tests check what the matrix literal guarantees by construction: the diagonal is exactly zero, $F_{\mu\nu}+F_{\nu\mu}$ is below $10^{-12}$, and reading $\mathbf E$ and $\mathbf B$ back out returns the inputs. There is one other project of mine that touches the same ground, [faraday](/faraday). The two repositories share no code and neither names the other. They share one idea: threshold a field, then summarise the thresholded set topologically. ## What it can do It generates three families of field, each in all three languages under the same names: `toroidal_pulse`, `braided_pair` and `boundary_sheaf`. Each family is a function of a point $(x,y)\in[-1,1]^2$, a normalised time $t\in[0,1]$, and three dials the code calls coupling $\kappa$, phase $\varphi$ and separation $s$. Grid size $N$ and the number of time steps $T$ are the other two inputs. The Rust core clamps $N$ to $[8,96]$, $T$ to $[1,128]$, $\kappa$ to $[0,1]$ and $s$ to $[0.05,0.9]$; the phase is not clamped. The **toroidal pulse** puts a ring of radius $0.34+0.28s$ on the grid. The design document describes it as an electric component that "flows tangentially around the ring while the magnetic component is radial with a phase-shifted axial component." The **braided pair** is two Gaussian lobes on opposite sides of the origin, rotating on a circle of radius $s$ as $t$ advances. The **boundary sheaf** places two Gaussian sources at $x=\pm s$ and a narrow band along $y=0$ that carries the out-of-plane components. The app is where most people will meet it. On the hosted page you pick a scenario and move six sliders: coupling (0.05 to 1), phase (0 to 1), separation (0.08 to 0.9), grid (16 to 72 in steps of 2), number of frames (4 to 48) and the current frame. The app recomputes the whole dataset in the browser whenever the configuration changes and steps through the frames every 280 ms while playing. It draws the energy density as a heat map with short strokes along $\mathbf E\times\mathbf B$, shows the Faraday tensor at the sample with the highest energy in the current frame, the two invariants at that sample, and the active-graph counts, and it exports the current dataset as JSON. Its default configuration is the toroidal pulse at $N=44$, $T=16$, $\kappa=0.72$, $\varphi=0.28$, $s=0.55$. **Figure 2.** The scenario generator run on your settings. Pick a family and set coupling, phase, separation, grid size, number of frames and frame (Play steps through them); the surface height and shade are the energy density u at each sample, cones show the in-plane direction of E×B, and a translucent plane sits at the repository's threshold τ = 0.38 × max u on the final frame, with the samples above it tinted. Click or use the arrow keys to place a probe; the matrix shows the Faraday tensor there, and the readout gives I₁, I₂, the mean energy, the coupling index, the two Betti proxies and the three residuals, exactly as the repository's code computes them. Time is normalised, and E×B carries no 1/μ₀ factor, so the arrows are a direction and relative size, not a flux in watts per square metre. Colour key: seq: energy density u; pos: E×B direction; parameter: threshold τ and the samples above it; neg: E cells of the tensor at the probe; structure: B cells of the tensor at the probe. The Python package is the reproducible path. `python -m em_field_sim --kind toroidal_pulse --grid 24 --steps 8 --out file.json` calls the same simulation and writes the full dataset, every sample of every frame plus a summary block, as indented JSON. The Rust crate exposes the same function, `simulate_scenario`, behind a `std`/`alloc`/`no_std` feature split, and its only dependency is `libm`. The repository commits one dataset, `web/public/data/sample-scenario.json`. Its configuration block says toroidal pulse, grid 8, one step, coupling 0.72, phase 0.28, separation 0.55. With one step the only frame is $t=0$, and the 8×8 grid has 64 samples. These are its summary numbers, copied from the file: **Measured: 0.0918** energy_total: mean energy density over the 64 samples of the committed fixture (toroidal pulse, 8×8, one frame). Control: none in the repository; no other committed dataset to compare against. n = 64 samples (N²·T = 8²·1). Source: web/public/data/sample-scenario.json:1555 @ f664a02 (exact value 0.09178557335741643). **Measured: β₀ 4 · β₁ 0** Betti proxies of the active set in the same fixture: 8 active vertices, 4 active edges, threshold 0.2435. Control: threshold τ = 0.38 × max u on the final frame; no other threshold is computed. Source: web/public/data/sample-scenario.json:1563-1567 @ f664a02. **Measured: 0.987** faraday_curl in the same fixture; divergence_e is 0.314 and divergence_b is 0.714. Control: none in the repository; the design document calls these educational diagnostics, not a solver guarantee. n = mean over the 36 interior samples, (N−2)² with N = 8. Source: web/public/data/sample-scenario.json:1558-1560 @ f664a02. The design document says the toroidal pulse "produces a stable active-cycle topology"; the committed fixture of that scenario reports four components and no cycle, because at 8×8 the ring does not survive thresholding as a ring. The fixture's configuration (grid 8, steps 1) also differs from the README's export command (`--grid 24 --steps 8`), and the app computes its own dataset in the browser rather than reading the file. ## How it was made The pipeline is the same in all three languages, function for function. In Rust it is `simulate_scenario` in `crates/em-field-sim/src/lib.rs`: clamp the configuration; for each step $k$ set $t_k=k/\max(T-1,1)$ and sample a frame; in each frame map row and column indices to $(x,y)$, call the scenario's field function, build the tensor, and record $\mathbf E\times\mathbf B$, the energy density and the invariants; accumulate energy and the coupling score over every frame; then compute the topology and the residuals on the final frame only, and divide the sums by $N^2T$. Python mirrors this in `simulate_scenario`, `_sample_frame`, `_topology_summary` and `_residual_summary`; TypeScript in `simulateScenario`, `sampleFrame`, `topologySummary` and `residualSummary`. ### The three field families With $\theta=2\pi(t+\varphi)$, a ring radius $R_0=0.34+0.28s$, $r=\max(\sqrt{x^2+y^2},10^{-6})$, unit vectors $\hat e_\theta=(-y,x,0)/r$ and $\hat e_r=(x,y,0)/r$, and the envelope $w=e^{-(r-R_0)^2/0.018}\,e^{-0.28(x^2+y^2)}$, the code writes the three families as: $$ \begin{aligned} \text{toroidal:}\quad &\mathbf E=w\,(0.72+0.28\sin\theta)\,\hat e_\theta+\bigl(0,0,\,w\kappa\cos(\theta+r)\bigr), &&\mathbf B=w\kappa\,\hat e_r+\bigl(0,0,\,0.25\,w\sin(\theta-r)\bigr)\\ \text{braided:}\quad &\mathbf E=\bigl((x-c_{1x})g_1-(x-c_{2x})g_2,\ (y-c_{1y})g_1-(y-c_{2y})g_2,\ \kappa(g_1+g_2)\sin\theta\bigr), &&\mathbf B=\bigl(-(y-c_{1y})g_1-(y-c_{2y})g_2,\ (x-c_{1x})g_1+(x-c_{2x})g_2,\ \kappa(g_1-g_2)\cos\theta\bigr)\\ \text{sheaf:}\quad &\mathbf E=\bigl(L-R,\ 0.25\,b\sin\theta,\ \kappa\,b\cos(\theta+x)\bigr), &&\mathbf B=\bigl(-0.2\,b\cos\theta,\ \kappa(L+R),\ b\sin(\theta+y)\bigr) \end{aligned} $$ **Figure 3.** Each field family, one component at a time. Six maps show E_x, E_y, E_z and B_x, B_y, B_z at the chosen time t for the chosen family; the E maps share one symmetric colour scale and the B maps share another, so the relative sizes are real. A toggle switches to the terms that build them: the ring envelope w and the four amplitude terms for the toroidal pulse, the lobes g₁ and g₂ and their moving centres for the braided pair, and L, R and the band b for the boundary sheaf. The strip below traces the six components at a probe point (click a map or use the arrow keys to move it) over t from 0 to 1 and marks that t = 1 repeats t = 0, because every family depends on t only through θ = 2π(t + φ). Colour key: neg: E components (signed: light for negative, saturated for positive); structure: B components (signed); seq: envelope, lobe and amplitude terms, each on its own range; parameter: current time t. In the braided pair the lobe centres are $\mathbf c_1=s(\cos\theta,\sin\theta)$ and $\mathbf c_2=-\mathbf c_1$, with $g_i=e^{-\lvert(x,y)-\mathbf c_i\rvert^2/0.08}$. In the boundary sheaf $L=e^{-((x+s)^2+y^2)/0.12}$, $R=e^{-((x-s)^2+y^2)/0.12}$ and $b=e^{-y^2/0.025}$. The helper that computes these lobes is named `gaussian(x, y, sigma)`, but its third argument divides the squared radius directly; it is not $2\sigma^2$. (Note: So the braided lobe "width" 0.08 corresponds to a standard deviation of 0.2, and the sheaf's 0.12 to about 0.245.) One property of these formulas matters later. Every family depends on $t$ only through $2\pi(t+\varphi)$, so $t=0$ and $t=1$ give identical fields. With $T>1$ the time samples run from $t=0$ to $t=1$ inclusive, which means the last frame, the one the topology and residuals are computed on, is the same field as the first. ### The invariants From the tensor the code reads two scalars at each sample. The Rust field names are `magnetic_minus_electric` and `pseudoscalar`. When I checked the matrix against the standard boost, it behaves as the covariant tensor in signature $(+,-,-,-)$ with $c=1$; that reading is mine, not the repository's, and the contraction below depends on it: $$ I_1=\lvert\mathbf B\rvert^2-\frac{\lvert\mathbf E\rvert^2}{c^2},\qquad I_2=\mathbf E\cdot\mathbf B,\qquad \tfrac12F_{\mu\nu}F^{\mu\nu}=I_1\ \ \text{(derived; }\eta=\mathrm{diag}(1,-1,-1,-1),\ c=1\text{)} . $$ **Figure 4.** (derived). I₁ and I₂ are what survives a change of inertial frame. The boost is not in the repository; this figure supplies it to test the repository's definitions on the repository's own tensor. Set E, B and a velocity β (as a fraction of c); the two 3D views show the fields in the original frame S and in the boosted frame S′, computed as F′ = LᵀFL and checked against the textbook boost formulas, with the largest difference shown. On the right, the point (|E|², |B|²) slides along a 45° line as β grows, and the strip plots I₁ and I₂ against β as flat lines while |E| changes. A button finds the frame that removes B (when E·B = 0 and |E| > |B|) or E (when |B| > |E|), and another sweeps β from 0 to 0.99 along x. Colour key: neg: E; structure: B; proof: I₁ and I₂, unchanged to rounding; baseline: |E|, which changes with the frame, and the boost direction β drawn in S; parameter: boost speed β. The design document calls these "Lorentz-style field invariants", and the tests check their definitions: for $\mathbf E=(0.5,0.25,-0.75)$ and $\mathbf B=(0.2,-0.4,0.6)$, $I_1=0.56-0.875=-0.315$ and $I_2=-0.45$, to $10^{-12}$ in Rust and TypeScript. No test applies a boost; the claim that the two numbers are frame-independent rests on standard electromagnetism, which Figure 4 runs live. The scenario fields themselves are not Lorentz-covariant solutions of anything. Only the pointwise algebra is relativistic. ### Energy, flow and the coupling score $$ u=\tfrac12\bigl(\lvert\mathbf E\rvert^2+\lvert\mathbf B\rvert^2\bigr),\qquad \mathbf S=\mathbf E\times\mathbf B,\qquad \bar u=\frac{1}{N^2T}\sum_{k,p}u_{kp},\qquad \mathcal C=\frac{1}{N^2T}\sum_{k,p}\Bigl(\lvert\mathbf E\cdot\mathbf B\rvert+0.05\,\lvert\mathbf S\rvert\Bigr) $$ **Figure 5.** The two headline numbers are averages, and the coupling score is a sum of two terms. Three rows of maps show u, |E·B| and 0.05|E×B| for the chosen family at five coupling values (0, 0.25, 0.5, 0.75, 1), each row on one shared scale so brightening with κ is real. The plot below stacks the two terms of the coupling score against κ and draws the mean energy as a line; at κ = 0 the score is not zero, because the E×B term remains. A frame slider picks the frame the maps show. The plot and the readout use the repository's formulas over all frames, as the summary does (the plot at N = 32 and T = 8, the readout at N = 44 and T = 16), and a toggle adds per-scenario bars to the readout. Colour key: seq: energy density u and its mean (the plot line); pos: |E·B| and 0.05|E×B| (the two terms of the coupling score; the second hatched); parameter: coupling κ. The code carries no permittivity or permeability: $u$ is $\tfrac12(E^2+B^2)$ in units where $\varepsilon=\mu=c=1$, and $\mathbf S$ has no $1/\mu_0$. The summary calls $\bar u$ `energy_total` although it is a mean, and the app labels it "Energy density". $\mathcal C$ is a score I built by hand: the weight 0.05 is a literal in the code, and the score has no physical unit and no baseline in the repository. Both averages run over every frame, unlike the topology and residuals. ### The active field graph > **Definition: Active field graph.** > > On the final frame, let $\tau=0.38\max_p u_p$. The active set $A$ is the set of grid samples with $u_p\ge\tau$. Its graph has one vertex per active sample and one edge per pair of active samples that are 4-neighbours (left, right, up, down). $\beta_0$ is the number of connected components of that graph and $\beta_1$ is reported as edges plus components minus vertices. The design document says "exceeds an adaptive threshold"; the code uses $\ge$ and the fixed fraction 0.38. In Rust the edges are counted once each by looking right and down from every active sample, the components by an explicit-stack depth-first search, and $\beta_1$ with saturating arithmetic; Python and TypeScript clamp at zero with `max`. Here is that function's core, verbatim at the pinned commit: ```rust // crates/em-field-sim/src/lib.rs:351-375 @ f664a02 (excerpt) let threshold = max_energy * 0.38; // ... let vertices = active.iter().filter(|&&is_active| is_active).count(); let components = count_components(&active, grid); let betti_1 = edges.saturating_add(components).saturating_sub(vertices); ``` The repository is clear about what this is: "This is a graph-cycle Betti proxy, not a full Vietoris-Rips persistence pass." $E_a+\beta_0-V$ is the cycle rank of the graph, the number of independent loops in it. When I checked the code, the gap between that and the number of holes in the active region turned out to have an exact form. Every 2×2 block of active samples closes a unit square, and each closed square is a loop in the graph that encloses nothing. If $F$ counts those blocks, the cubical complex with the squares filled in has Euler characteristic $V-E_a+F$, and in the plane its first Betti number is $\beta_0$ minus that: $$ \beta_1^{\text{repo}}=E_a+\beta_0-V,\qquad \beta_1^{\square}=\beta_0-\bigl(V-E_a+F\bigr),\qquad \beta_1^{\text{repo}}=\beta_1^{\square}+F\ \ \text{(derived)} . $$ **Figure 6.** (derived bookkeeping). A threshold makes a graph, and the graph's cycle count includes every filled square. On the final frame of the chosen family, seen from above, the active samples are dots, their 4-neighbour links are edges, every 2×2 block of active samples is shaded, and each enclosed inactive region (a true hole) is ringed. Chips give V, E, F and C, then the repository's β₁ as holes plus F and the hole count of the filled-in complex, which is checked live against an independent count of enclosed inactive regions. Move the threshold fraction (the repository's 0.38 is marked) and the grid size to watch F grow with area while the hole count stays put. A solid 3×3 block is a one-click test: 9 vertices, 12 edges, 4 squares, proxy 4, holes 0. The identity is Euler-characteristic algebra, not a statement the repository makes. Colour key: seq: energy density u (the surface); structure: active samples and their 4-neighbour edges; withdrawn: filled 2×2 blocks, each adding 1 to the repository's β₁; proof: enclosed holes (the β₁ of the filled-in complex); parameter: threshold fraction. The numbers make the point. When I checked the code with the app's default configuration (toroidal pulse, $N=44$, $T=16$, $\kappa=0.72$, $\varphi=0.28$, $s=0.55$), the final frame has 268 active samples, 448 active edges, 180 filled 2×2 blocks and one component. The repository's $\beta_1$ is 181. The ring has one hole. So at the app's default settings the Python port reports $\beta_1=181$ for a field whose active region is an annulus, and the app displays the same count from its TypeScript port (expected, not observed; see the last section). The decomposition held in all 162 cases I ran (three families, $N\in\{16,24,44\}$, three phases, six threshold fractions): the repository's number always equalled the independent hole count plus $F$. The proxy is not wrong for what its tests ask, which is $\beta_1\ge1$; it counts loops in a graph, not holes in a region. ### The residuals With $\Delta=2/(N-1)$, central differences on the interior samples, row index $i$ along $y$ and column index $j$ along $x$, the three numbers are means over the $(N-2)^2$ interior samples of the final frame. A fourth expression, which the repository does not compute, is the one Faraday's law actually constrains: $$ \begin{aligned} \partial_xf\big|_{ij}&=\frac{f_{i,j+1}-f_{i,j-1}}{2\Delta},\qquad \partial_yf\big|_{ij}=\frac{f_{i+1,j}-f_{i-1,j}}{2\Delta},\\ \texttt{divergence\_e}&=\frac{1}{(N-2)^2}\sum_{i,j=1}^{N-2}\frac{\lvert\partial_xE_x+\partial_yE_y\rvert}{1+u_{ij}},\qquad \texttt{divergence\_b}=\frac{1}{(N-2)^2}\sum_{i,j=1}^{N-2}\frac{\lvert\partial_xB_x+\partial_yB_y\rvert}{1+u_{ij}},\\ \texttt{faraday\_curl}&=\frac{1}{(N-2)^2}\sum_{i,j=1}^{N-2}\frac{\lvert\partial_xE_y-\partial_yE_x\rvert}{1+u_{ij}},\qquad R(\sigma)=\frac{\lVert\nabla\times\mathbf E+\sigma\,\partial_t\mathbf B\rVert}{\lVert\nabla\times\mathbf E\rVert}\ \ \text{(derived)} . \end{aligned} $$ **Figure 7.** (derived control). What the three residuals read on a field that does satisfy Maxwell's equations. Left: the repository's scenario on its final frame, with maps of B_z, of |(∇×E)\_z|/(1+u) (what faraday_curl averages), and of the full Faraday residual |∇×E + σ ∂B/∂t| using the time derivative of the repository's own field function. Right: a 2D TEz Yee finite-difference time-domain cavity (in-plane E, out-of-plane B, perfectly conducting walls), which is not in the repository and serves as the control; run it to watch a pulse expand and reflect. A table puts the three repository metrics side by side for both fields, with R at σ = 1 and at the best-fitting σ. The control's faraday_curl is far from zero although its full residual is zero to rounding, and the scenario's full residual stays near 1 for every σ. Colour key: structure: B_z (signed); pos: |(∇×E)\_z|/(1+u), the repository's faraday_curl integrand; seq: full Faraday residual; baseline: TEz Yee control; withdrawn: residual near 1 at every time scale; proof: residual zero to rounding. The code computes these lines directly: ```rust // crates/em-field-sim/src/lib.rs:441-446 @ f664a02 let curl_z = (right.electric.y - left.electric.y - up.electric.x + down.electric.x) / (2.0 * dx); let scale = 1.0 + samples[idx].energy_density; div_e += ((d_ex_dx + d_ey_dy) / scale).abs(); div_b += ((d_bx_dx + d_by_dy) / scale).abs(); curl_proxy += (curl_z / scale).abs(); ``` The number named `faraday_curl` is the $z$-component of the spatial curl of $\mathbf E$, divided by $1+u$ and averaged. There is no $\partial\mathbf B/\partial t$ term anywhere in the residual code, so it does not measure Faraday's law; a travelling wave that satisfies the law exactly has a non-zero curl of $\mathbf E$. When I checked the code against a Yee time-domain control, the control's `faraday_curl` came out between 0.41 and 0.54 at four time steps, while its full residual was below $10^{-6}$ at all four, and its discrete energy read 0.0628319 at steps 20, 40 and 80. I then asked the opposite question of the scenarios: is there any time scale $\sigma$ at which $\nabla\times\mathbf E=-\sigma\,\partial_t\mathbf B$ holds for the field I wrote? The best-fit relative residual was at least 0.998 for all three families at $N=44$, $t=0.5$ (toroidal 0.999, braided 1.000, sheaf 1.000), and for the toroidal pulse the best-fitting $\sigma$ even had the wrong sign. At that grid and time, none of the synthetic fields satisfies Faraday's law at any time scale. That is consistent with what the repository says they are, "finite-difference educational residuals", but it means a smaller `faraday_curl` is not a more Maxwell-like field. ## What's new in it Two usual approaches sit on either side of this toolkit. The usual way to teach electromagnetism interactively is to draw fields: field lines, arrow grids, coloured potentials, or a running wave simulation. [PhET's Faraday's Law](https://phet.colorado.edu/en/simulations/faradays-law) moves a magnet through a coil; [Paul Falstad's 2D electromagnetic wave applet](https://www.falstad.com/emwave2/) solves a time-domain wave problem in the browser and shows $\mathbf E$ and $\mathbf B$ directly. By their published descriptions, neither puts the Faraday tensor on the screen. Here the tensor is the primitive: every sample builds $F_{\mu\nu}$, the invariants are read from it, and the app shows the 4×4 matrix at the peak-energy sample next to $B^2-E^2$ and $\mathbf E\cdot\mathbf B$. The cost is that the fields are written down rather than solved for. The usual way to summarise a scalar field topologically is a filtration: sweep a level through the whole range of the field and record when components and holes are born and die, as persistent homology or a Morse–Smale complex does. This toolkit takes one level, 0.38 of the maximum energy, and counts the graph at that level. The design document says the choice is deliberate: "the browser can recompute the field and topology instantly as users drag controls." Figure 8 puts the two side by side on the same field. Its filtration is derived, not in the repository: $$ A_f=\{\,p:\ u_p\ge f\cdot\max_q u_q\,\},\qquad f\in[0,1],\qquad A_{f'}\subseteq A_f\ \text{ whenever } f'\ge f . $$ **Figure 8.** (derived baseline). One fixed threshold against the whole superlevel filtration of the same energy map. Cells are added in decreasing order of u with a union-find, so a single pass gives, for every threshold fraction f from 0 to 1, the vertex, edge, filled-square and component counts, the repository's β₁ proxy, and the hole count of the filled-in complex. A cursor sets f for the readout and lights the bars alive at that level, a toggle switches the count axis between log and linear, and a dashed vertical line marks the repository's f = 0.38; a band on the axis marks the range of f over which the active region is one component with exactly one hole. Under the plot, the H₀ barcode of the filtration shows one bar per component, born at a local peak of u and ending where it merges into an older one. The filtration is the standard method the design document points to for higher-fidelity work, computed here as a baseline; it is not the repository's code. Colour key: proof: β₁ of the filled-in complex (holes); structure: the repository's β₁ proxy; withdrawn: filled 2×2 blocks F; baseline: component count C and the H₀ barcode; parameter: threshold fraction f; the repository's 0.38 is marked. What is different, then, is the combination and its speed, not any one piece. The same three files of arithmetic produce the field, the tensor, the invariants, the energy graph and the residuals, so one configuration gives one dataset whichever language you run, and the browser can redo the whole thing between two slider events. The single threshold is the price of that speed; the filtration is what the design document leaves to "the existing EMFS persistent-homology stack", which this repository does not contain. ## What no one else built I looked for the closest prior work in four directions: browser simulations for teaching electromagnetism, research field solvers, visualisations of relativistic field transformations, and topological summaries of scalar fields. Each of the pieces here has a predecessor; what I can claim is narrower than any of them. **Teaching simulations.** [PhET's Faraday's Law](https://phet.colorado.edu/en/simulations/faradays-law) and [Falstad's 2D wave applet](https://www.falstad.com/emwave2/) are browser tools built for classrooms. Falstad's applet is a real solver: it covers reflection at conductors, dielectric boundaries, dipole radiation and waveguides, which overlaps what my design document files under "future layers" (material models and boundary-condition solvers). MIT's TEAL project went further on energy flow; Belcher and Koleci's [animated textures](https://arxiv.org/abs/0802.4034) show field motion and energy transport with a second velocity field. None of these, as they are described, shows $F_{\mu\nu}$ or the two invariants or exports a dataset; my toolkit does both and solves nothing. **Research solvers.** [Meep](https://meep.readthedocs.io/) is the standard open-source finite-difference time-domain package, with materials, sources and boundary conditions. It produces fields that satisfy a discretised Maxwell's equations. My toolkit does not compete with it, and Figure 7's control exists because a Yee scheme like Meep's is the right reference for what a "Maxwell residual" should read. **Relativistic field visualisation.** [Seeing through the light cone](https://arxiv.org/abs/2505.20596) extends a past-light-cone visualisation method to electromagnetic fields, and is the closest work I found to showing how $\mathbf E$ and $\mathbf B$ transform between frames. My repository does not boost anything; it computes the invariants at fixed frames. Figure 4 adds the boost, and is marked derived for that reason. **Topology of scalar fields.** The [Topology ToolKit](https://topology-tool-kit.github.io/) ([Tierny et al.](https://arxiv.org/abs/1805.09110)) computes persistence diagrams, merge trees and Morse–Smale complexes of scalar fields on regular grids. [GUDHI's cubical complexes](https://gudhi.inria.fr/python/latest/cubical_complex_user.html) and [Cubical Ripser](https://arxiv.org/abs/2005.12692) compute persistent homology of grid data with the squares filled in, which is exactly the correction Figure 6 shows my proxy lacks. Persistence itself goes back to [Edelsbrunner, Letscher and Zomorodian](https://doi.org/10.1007/s00454-002-2885-2). Applied to electromagnetism, Gross and Kotiuga's [Electromagnetic Theory and Computation: A Topological Approach](https://library.slmath.org/books/Book48/desc.html) uses cohomology for boundary-value problems, and a 2024 APS abstract by [Bohlsen, Robins and Hole](https://archive.aps.org/dpp/2024/tp12/49) applies persistent homology to magnetic field-line orbits. Every one of these does more topology than my single-threshold graph. What survives the comparison is one specific arrangement. In a single browser loop, the same code builds the Faraday tensor at every sample, reads both invariants, thresholds the energy density and summarises the thresholded set as a graph, and computes three finite-difference diagnostics, and the same pipeline exists in Rust, Python and TypeScript (the JSON export is in the Python command line and the app; the Rust crate has none). I did not find another tool that puts a Faraday-tensor readout and a topological summary of the same field's energy on one screen. That is the extent of the claim: an arrangement for teaching, not a method. **Figure 9.** The same mathematics in three languages, with the repository's assertions run live. A table re-runs every assertion from the three test files (antisymmetry, field round trip, invariant definitions, and each file's scenario test) on the figure's TypeScript port with the repository's own inputs, compares the port's output for the committed fixture configuration with the fixture's stored digits, and fuzzes 1,000 seeded random field pairs (a seed selector and a re-run button repeat the fuzz). A 3×3 coverage matrix shows which scenario family each language's tests exercise: the Rust and TypeScript tests use only the toroidal pulse, the Python test only the braided pair, no test uses the boundary sheaf, and no test compares one language's output with another's. The Rust crate itself is not run here; its assertions are re-run on the port. Colour key: proof: an assertion that holds; withdrawn: an assertion that fails, or a family a language's tests do not run; measured: a test file runs this scenario family (coverage matrix); baseline: cross-language comparison: none in the repository. ## Limitations **It is not a solver.** The fields are closed-form patterns, not solutions. The repository says the residuals are "not a substitute for a production Maxwell solver with boundary conditions, material models, and stability proofs", and lists material models, anisotropic media, boundary-condition solvers and GPU kernels as future layers. Figure 7 quantifies the gap: at $N=44$ and $t=0.5$, no time scale makes any of the three families satisfy Faraday's law. **The residual named `faraday_curl` is not Faraday's law.** It is the spatial curl of $\mathbf E$ with no $\partial\mathbf B/\partial t$ term. It is non-zero for any travelling wave and small for any nearly curl-free $\mathbf E$, so its size says nothing about how Maxwell-like a field is. **The $\beta_1$ proxy counts squares.** It equals the number of holes plus the number of filled 2×2 blocks. The repository documents it as a proxy and its tests ask only $\beta_1\ge1$, which any blob of four active samples satisfies. **Only the final frame is analysed, and with $T>1$ it is the first frame again.** Topology and residuals ignore every frame but the last, and because the fields are periodic in $t$ with period 1, the last frame repeats $t=0$. The energy and coupling averages, by contrast, count that frame twice. **The residuals and counts do not converge to the quantities their names suggest.** When I checked the code across the clamp range of grid sizes, the toroidal `divergence_e` fell like $\Delta^2$ (0.0149 at $N=44$, 0.0030 at $N=96$) because its in-plane $\mathbf E$ is divergence-free by construction, but toroidal `divergence_b` settled near 0.88, braided `divergence_e` near 0.159, and the boundary sheaf's three numbers near 0.45, 0.38 and 0.51. Those are non-zero limits, properties of the formulas rather than of the grid. The toroidal proxy grew from 33 at $N=24$ to 181 at $N=44$ and 1,121 at $N=96$, while the hole count stayed at 1. **Figure 10.** What the diagnostics do as the grid refines from 8 to 96. One row per family, three plots per row, all on the final frame and sharing the grid-size axis: the three residuals, the topology counts (components, the repository's β₁ proxy, and the hole count of the filled-in complex), and the number of active samples. The sweep runs the repository's own formulas and fills left to right over about 3 seconds. A dashed line marks the Rust test's bound of 0.25 on divergence_e; in the toroidal row a solid square marks the committed fixture's divergence_e at N = 8; residuals that settle at a non-zero level are labelled as limits. Colour key: neg: divergence_e; structure: divergence_b; pos: faraday_curl; proof: hole count of the filled-in complex; withdrawn: the repository's β₁ proxy; ink-2: components and active samples; baseline: Rust test bound 0.25; measured: committed fixture divergence_e at N = 8; parameter: grid size N. **What failed.** - Withdrawn: The toroidal pulse produces a stable active-cycle topology. Killed by: The committed fixture of that scenario (8×8, one frame) reports β₀ 4, β₁ 0; at larger grids the reported β₁ is mostly filled squares (181 at N = 44 against one hole, from my check of the code).. - Withdrawn: faraday_curl is a Maxwell residual for Faraday's law. Killed by: lib.rs:441-446 computes only (∇×E)\_z/(1+u); no ∂B/∂t term exists in residual_summary.. - Withdrawn: The boundary sheaf checks whether neighbouring patch summaries remain compatible. Killed by: The code defines only the field function (lib.rs:325-341); no patch, restriction or compatibility computation exists in any of the three languages.. - Withdrawn: sample-scenario.json is the output of the README's export command. Killed by: The README command uses grid 24, steps 8; the file's configuration says grid 8, steps 1.. - Withdrawn: Pushes to main deploy the Pages site. Killed by: The workflow triggers on master (pages.yml:4-5).. - Withdrawn: The Faraday tensor is implemented and tested in Rust, Python and TypeScript. Killed by: Implemented, yes; but no test compares the three outputs, and each test file imports only its own module.. **Where the documents and code differ.** The design document lists `npm audit` among the checks, but the Pages workflow runs `npm ci`, `npm test` and `npm run build`. It says vertices "exceed" the threshold where the code uses $\ge$. The README calls the simulator "high-performance", and the repository has no benchmark (the design document lists benchmark fixtures as future work). The app's "Maxwell residual" tile shows only `divergence_e`. **Test coverage is thin and one family is untested.** Rust tests the toroidal pulse, Python the braided pair, TypeScript the toroidal pulse; nothing tests the boundary sheaf. The scenario assertions are loose: $\beta_0\ge1$, $\beta_1\ge1$, positive energy, and `divergence_e` below 0.25 (Rust) or 0.35 (Python). **The three implementations can disagree on non-integer input.** TypeScript rounds `grid` and `steps` with `Math.round`, Python truncates with `int()`, and Rust takes an unsigned integer, so a grid of 24.6 gives 25 in one language and 24 in another. Python also falls back silently to the toroidal pulse for any unknown `kind`; only its command-line parser restricts the choices. **What I did not do.** There are no benchmarks of any kind in the repository, so there is no speed claim to report. There is no Lean or other formal proof; the only guarantees are the antisymmetry and round-trip tests, which hold by construction of the matrix. I did not build or run the Rust crate or the TypeScript module for this essay. The numbers marked "when I checked the code" come from running the Python module's own functions, which reproduce the committed fixture digit for digit; Rust and TypeScript agreement with them is expected from the three sources mirroring one another function for function (apart from the rounding rule above), not observed. The time-domain control, the boost, the filled-square identity and the filtration are mathematics I added to test the repository, and each is labelled derived. ## Read more [View the project](https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator) · [Source on GitHub](https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator) The hosted app is at [teerth.dev/Electromagnetic-Field-Data-Simulator](https://teerth.dev/Electromagnetic-Field-Data-Simulator/) and the source at [github.com/teerthsharma/Electromagnetic-Field-Data-Simulator](https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator). Version 0.1.0 is archived on Zenodo as [10.5281/zenodo.21997896](https://doi.org/10.5281/zenodo.21997896). The files that carry the substance, all at commit `f664a02`: - [`docs/EM_FIELD_DATA_SIMULATOR.md`](https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator/blob/f664a02ac981f3f4bb785d7015904e1c4f240386/docs/EM_FIELD_DATA_SIMULATOR.md): the design document, scenario descriptions and stated limits. - [`crates/em-field-sim/src/lib.rs`](https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator/blob/f664a02ac981f3f4bb785d7015904e1c4f240386/crates/em-field-sim/src/lib.rs): the Rust reference core; [`crates/em-field-sim/tests/faraday_contract.rs`](https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator/blob/f664a02ac981f3f4bb785d7015904e1c4f240386/crates/em-field-sim/tests/faraday_contract.rs): its tests. - [`python/em_field_sim/core.py`](https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator/blob/f664a02ac981f3f4bb785d7015904e1c4f240386/python/em_field_sim/core.py): the Python mirror and the command-line exporter. - [`web/src/sim/faraday.ts`](https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator/blob/f664a02ac981f3f4bb785d7015904e1c4f240386/web/src/sim/faraday.ts) and [`web/src/App.tsx`](https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator/blob/f664a02ac981f3f4bb785d7015904e1c4f240386/web/src/App.tsx): the TypeScript port and the app. - [`web/public/data/sample-scenario.json`](https://github.com/teerthsharma/Electromagnetic-Field-Data-Simulator/blob/f664a02ac981f3f4bb785d7015904e1c4f240386/web/public/data/sample-scenario.json): the committed 8×8 fixture. Related essays on this site: [faraday](/faraday), which shares the idea of thresholding a field and summarising the thresholded set topologically, and [topological-ml-toolkit](/topological-ml-toolkit), a separate library of mine for topology in machine-learning pipelines, including persistent homology.