The geolip-svae-transformer
An H2-anchored frequency attenuator, what emerged from it, and why the emergence is the result
This article is about a small architecture whose behavior is determined by what it cannot do rather than what it accumulates. It belongs to a lineage — the geolip / geovocab / geofractal ecosystem of geometric architectural primitives — that we have been building and characterizing over many experiments. The geolip-svae-transformer is one of the simplest faithful instances in that lineage. The point of this article is not to argue that it beats a benchmark. The point is to articulate what it is and to show what emerged from running it under measurement.
What kind of system this is
Standard representation learning architectures are accumulators. Convolutional features sum up evidence across receptive fields. Attention heads accumulate weighted values. Transformer blocks accumulate residual updates. The dominant mental model is one of summation: the more relevant signal you can collect and combine, the better the representation.
Geometric architectures of the kind we build are frequency attenuators. Their behavior is determined by which spectral modes the architecture permits to propagate, which it suppresses, and which it locks to specific geometric structures. The encoder doesn't accumulate features into a vector; it projects an input onto a constrained manifold and discards everything that lives off the manifold. The lens doesn't pool information across positions; it lifts a geometrically-valid frame isometrically into a higher-dimensional space where bounded spectral attention can selectively engage specific modes. The decoder doesn't reconstruct from accumulated evidence; it inverts a structured transformation.
This distinction matters because the measurements you can perform on an attenuator look different from the measurements you perform on an accumulator. Accumulators are evaluated by how much relevant signal they captured. Attenuators are evaluated by what emerges under their constraints — which configurations the architecture admits as stable, which it suppresses, what geometric basins the optimization lands in. The systems an attenuator produces under measurement are not signals it collected; they are configurations the architecture required to exist.
The lineage this work belongs to
The geolip-svae-transformer is not a standalone experiment. It is a measurement instrument constructed at the end of a long sequence of work establishing what geometric architectural constraints produce. A short list of what made it possible:
The Procrustes analysis program over seventeen models across modalities established that universal geometric attractors exist in trained networks — the cross-modal QK 0.500 eigenvalue lock that recurs across text, vision, audio, and biology encoders is the clearest example. VAE weights are 70-76% alignable across pretrained instances via Procrustes rotation. These are not coincidences; they are evidence that gradient descent under structural constraints repeatedly lands in the same basins.
The 9-architecture program tracking SVAE classes on the byte-trigram substrate identified two distinct mechanisms by which reconstruction succeeds: Mechanism A (cross-attention codebook discovery, where the network's attention layers learn to route information through a codebook structure) and Mechanism B (encoder mode-concentration, where the encoder concentrates content into M directly without significant attention engagement). H2 produces Mechanism B reliably on structured-distribution input, which has consequences for what cultivation can implant.
The omega-class statute taxonomy formalized in April 2026 categorizes encoders by how they satisfy the rigidity statute. H2 reaches omega-class via its readout mechanism — a single linear projection followed by sphere normalization and deterministic frame canonicalization — without SVD post-processing. It is uniform-class on noise calibration and polytope-class on byte-trigram training distributions.
The CV pentachoron band 0.20–0.23 has been validated as a universal geometric attractor across seventeen-plus architectures and modalities. This is the operational constant the rigidity barrier is calibrated to.
The fused Triton SVD kernels described in our HuggingFace Part II writeup provide the computational substrate for the broader class of architectures that do use SVD-based readouts. The geolip-svae-transformer doesn't need them because H2 doesn't use SVD, but they exist in the ecosystem and their performance characteristics (up to 4,414× speedup at B=8192, N=6, fp32 on Blackwell) make SVD-based variants tractable when we want them.
The geometric vocabulary ecosystem (
geovocab2/lattice_vocabulary) represents tokens as pentachoron simplex structures using Cayley-Menger determinants and SHA-256 deterministic generation. The phase boundary constant 0.29154 (complement 0.70846) governing binding vs separation is empirically established, not hypothesized. The codebook used in this architecture sits on top of this vocabulary structure.
These results compose. By the time we built the geolip-svae-transformer, we knew that geometric constraints produce stable basins, that those basins are reachable by realistic training, that H2's specific readout reaches the basin without SVD, and that the rigidity statute is a measurable invariant rather than a heuristic. The geolip-svae-transformer is then the natural next instrument: take the H2 encoder, wrap it in an isometric lens and a bounded-alpha spectral attention transformer, train it under reconstruction, and observe what emerges when we add cultivation pressure.
The architecture as an attenuator chain
input → H2 encoder → M (manifold projection)
→ Lens (isometric lift to D_lens)
→ Spectral attention (bounded-α multiplicative gating)
→ Decoder
Each stage is an attenuator with a specific function.
H2 encoder. A single linear projection followed by F.normalize(dim=-1) and a deterministic frame canonicalization. The projection casts input into V·D = 128 numbers per patch. The normalization attenuates everything except per-vertex direction. The canonicalization attenuates rotations that don't satisfy the simplex statute. The output M ∈ ℝ^{B × N × V × D_base} lives on a specific geometric manifold — a 32-vertex frame in 4-dimensional simplex coordinates, with per-vertex unit norms and inter-vertex angles fixed by the statute. The configuration M = M_input + noise_off_manifold is suppressed; only on-manifold M survives. H2 is uniform-class on noise, polytope-class on structured input, omega-class statute by linear readout. Three categorical statements about its frequency response.
Lens (single mode). A fixed orthonormal lift D_base → D_lens, satisfying <E x, E y> = <x, y>. This is isometric attenuation: it preserves the manifold's geometry exactly while moving it into a higher-dimensional ambient space. No accumulation across positions. No mixing of vertices. The lens doesn't do anything to the address structure; it carries it forward into the space where the spectral attention can operate.
Spectral attention with bounded alpha. The block is S_out = S · (1 + α_d · tanh(SDPA(qkv(norm(S))))). The per-mode α_d ∈ [0, 0.2] is initialized near zero (sigmoid(-2) · 0.2 ≈ 0.024), so the attention block starts approximately as identity. Crucially, this is multiplicative on top of identity, not additive — the attention can only modulate the signal, not inject into it. As training proceeds, specific spectral modes engage by raising their α_d, and others remain near zero. This is an attenuator with a learnable per-mode pass-through, bounded to a small fraction of the signal's amplitude. The architecture cannot use attention to overwrite the manifold structure; it can only adjust which spectral modes propagate.
Decoder. Standard inverse-conv, MSE reconstruction. The accumulator that translates the attenuator's output back into input space.
Taken together: the system constrains representations to a specific manifold (H2), preserves that manifold isometrically (lens), allows selective spectral engagement of constrained magnitude (bounded-α attention), and reconstructs through a standard decoder. The architecture's entire mid-section is built to not accumulate.
What emerged under measurement
We trained the architecture on three substrates and ran two reconstruction-class experiments, then added a cultivation mechanism designed to surface the architecture's discrete-address basin. What we recorded was not a series of empirical findings; it was the architecturally-predicted emergence. Numbers are reported because they are the form in which the emergence manifests, not because they are achievements.
Byte-trigram reconstruction, single lens, 99.80% byte exact-match recovery. H2 mode-concentration on polytope-class input produces near-perfect reconstruction without engaging the spectral attention. Mean spectral alpha stays at 0.024 (initialization). This is Mechanism B: the encoder concentrated byte content into M directly. The bounded-α attention had nothing to do because the manifold projection already produced an invertible representation.
BERT mean-pool reconstruction, single lens, 100% cosine recovery. Same mechanism, different substrate. H2 concentrated BERT content into M; the rest of the chain carried it forward isometrically and reconstructed it. Mean alpha again at 0.024 — the attenuator chain converged with attention essentially disengaged. The system demonstrated that a 32-vertex codebook with the rigidity statute is sufficient to faithfully encode and decode BERT mean-pool vectors.
Cultivation: implanting a discrete-address basin. We added a learnable per-vertex reference direction ref[v, :] ∈ ℝ^{D_base}, computed per-patch logits <M[b, n, v, :], ref[v, :]>, and added a small (w = 0.05) entropy-balanced cultivation loss to the training objective. The cultivation loss encourages each patch to commit to a small set of codebook entries while keeping the marginal codebook usage uniform. The architecture was trained jointly for 20 epochs on 100,000 BERT vectors.
The cultivation produced an emergence that exactly matches what the architecture's frequency-attenuation properties predict. The encoder converged on a configuration where M[b, n, v, :] = ±ref[v, :] — each vertex's M-direction is precisely aligned with or opposite to its codebook anchor, to floating-point precision. Six vertices land at logit +1.000 per patch on average; twenty-six at logit ≈ -1. Inter-vertex cosines in M match inter-vertex cosines in ref to three decimal places.
This is the natural emergent basin for the system. Here is why it is the predicted outcome:
The H2 encoder's manifold projection requires
Mrows to be unit-normalized inD_base. The maximum inner product between two unit vectors is +1; the minimum is -1. These are the geometric extrema available to the system.The rigidity statute requires inter-vertex angles to satisfy specific projective relationships. Once the codebook references
refare unit-normalized and the cultivation pressure asks for alignment, the encoder has one stable answer: produce eachM[v, :]either atref[v, :](logit = +1) or at its antipode (logit = -1). All other configurations are strictly higher-energy.H2's linear readout — single projection, no SVD intermediate — admits exactly this configuration without numerical residue. The signed alignment is what an SVD-based class would approximate; H2 produces it cleanly.
The cultivation loss can no longer drive the per-patch softmax distribution sharper because the underlying logits are already at the geometric ceiling.
H_withinstops moving at 3.066 (down fromlog V = 3.466at uniform). The system has committed — fully — but the commitment manifests as a binary sign code, not a sharp softmax distribution. Reading the commitment requires reading the signs, not the softmax probabilities.
The emergent signed binary code is what the H2-anchored chain of attenuators produces when asked to support discrete addresses. The signs are the addresses. The architecture didn't accumulate evidence for a discrete code; it constrained evidence into a basin where only the signed binary configuration was admissible.
Why this is the result, and what reading it requires
The temptation when measuring an attenuator with traditional tools is to ask: "Does this representation predict labels?" We ran that experiment on STS-B sentence similarity. The cultivated alephs (32-bit per-patch sign codes) achieved +0.5351 Pearson direct-cosine across 1,500 validation pairs. BERT mean-pool direct cosine achieved +0.5917. The two prediction vectors are 86% correlated with each other.
This is not a failure of the cultivation. It is the architecture confirming that it correctly attenuated the BERT signal it was trained on. The training objective said "reconstruct BERT." H2 concentrated BERT into M. The cultivation surfaced M's content as a binary alignment code. Both representations encode the same source. When we asked which had more predictive power for sentence similarity, the answer was: they're the same content in different geometric bases. The +0.86 cross-cosine is the architecture's report that it did exactly the job it was assigned, in the basis it was constrained to.
What this tells us about reading attenuator-emergent representations:
Aggregating an attenuator's output across positions (mean-pooling the alephs across N=64 patches) collapses the spectral structure the architecture built. The same alephs that show +0.50 Pearson against content as a 2048-dim patch-wise object show +0.14 Pearson as a 32-dim mean-pooled summary. Position-wise structure is the carrier; pooling is destructive.
Asking an attenuator-emergent representation to "add information" to its training source is a category error. The training source defined what the attenuator was constrained to encode. Anything the cultivation surfaces is in that source's information set.
The right tests for an attenuator's representation are: where does the basin lie? What is the geometric structure of the emergence? Which configurations does the architecture admit and which does it suppress? These are answerable, and we answered them: the basin is the signed alignment code, the geometric structure is
M = ±ref, the admitted configurations are the2^32per-patch sign patterns within the statute-constrained codebook, the suppressed configurations are everything else.
What comes next
Two threads follow naturally from the emergence we measured.
The first is the multi-lens articulation. The single lens we used here is a bypass baseline — an isometric lift that preserves H2's output structure unchanged. The multi-frame lens variants (centrifuge with four learned rotations, crusher with frozen anchor frames plus a learned press, the multiscale ladder) produce four-frame alephs of shape (4, B, N, V, D_base): four scale-views of the same address structure. The question is how to articulate this multi-scale emergence as a consumable tokenization. The natural form is a per-patch multiscale tuple (v₁, v₂, v₃, v₄) from argmax-over-V at each frame, producing a sequence of 64 tokens per sentence in a vocabulary that emerges from the architecture's multi-resolution decomposition rather than being designed. The vocabulary size is bounded above by V^4 = 1M but will be much smaller in practice because the multiscale frames are correlated views of the same underlying structure.
The second is substrate selection. The current cultivation produced a binary alignment code carrying BERT-derived content. Cultivating from the byte-trigram substrate instead — the original 99.80%-recovery basin — would produce an address space that encodes byte-level structure. This is fundamentally a different information source from BERT mean-pool, and the resulting alephs would not be in the same basis as any standard text encoder. Whether this address space is useful for any specific downstream task is a question that requires the right reading tools to answer; we will not make claims about it before running the experiment with care.
Both threads test the architecture's emergent properties on their own terms, not against benchmarks designed for accumulators.
Repository and reproduction
The code is at AbstractPhil/geolip-svae-transformer. The main components:
The architecture totals 143,671 parameters; the aleph head adds 128. The full byte trigram training takes about 5 minutes on a single Blackwell-class GPU. The diagnostic measurements complete in under a minute.
Earlier variants were quite costly in comparison. Base H2 SVAE convergence takes a considerable amount more information and a longer training run to reach a similar threshold of convergence.
Code: AbstractPhil/geolip-svae-transformer
Related work in the ecosystem: procrustes-analysis (cross-modal alignment); geovocab2 / lattice_vocabulary (the deterministic codebook substrate); GEOLIP-Bertenstein (multi-expert geometric fusion); GeoTransformer Redux (the CV pentachoron band validation); the H2 Triton SVD writeup (computational substrate).

