AbstractPhil
·
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
repliedto their post about 9 hours ago Mini-Beatrix-2s is cooking with full splat attention through and through. This model is still trigram, I did this to get a baseline because there's already a trigram model to compare to. This one should be done in a few days and ought to be substantially more intelligent than the first.
Specs are;
Around 220m params, 4096 context window, d1024 model size, splat 128, 1024, 1024, 1024, and so on, 3 experts per block, information banks for storage and retrieval, and a lot of technical knowhow between A to B.
Differences with V2;
Special tokens are implemented byte-directly, so the model will have no problem recognizing an array of special tokens such as DOC, EOF, and a multitude of others.
Suffice it to say, this model is bigger than the first at about 2x. Not just bigger though, estimated to be roughly 8x more intelligent based on the measures.
That being said, the actual model needs to be substantially larger to encompass the full space. The measured space is considerably larger through the small tests for stability, however the full 900m version runs at only around 8k tokens per second with an anchor count of 131,000 and a matching number of heads. This means the full train would require roughly 26 days on a rtx 6000 pro blackwell, which is substantially beyond the expectation curve.
So the smaller one will do for now until I can secure a bit of funding. In any case, the tokenizer system will be implemented on this version after a stable run completes. View all activity Organizations
view article Raising Beatrix: A Byte-Level Model's Measured Childhood
AbstractPhil
• view article Agreement, Anchors, Addresses: A Week of Geometric Training
AbstractPhil
• view article Geometric Memory FT4 — Distill Against a Consensus, Ship a Rotation
AbstractPhil
• • 1
published an article about 1 month ago view article The Loss Manifest: A Field History of Objective Functions, and What a Machine Can Actually Be Asked to Compute
published an article about 1 month ago view article Aleph Differentiation, Parts 3 & 3-D: Two Laws, Five Days, One Framework
published an article about 2 months ago view article The Aleph Moves Into a Pretrained Trunk: Relays, Registers, and the Two-Regime Dispatch Law
published an article about 2 months ago view article The Aleph Under Autoregressive Pressure: Bottleneck Priors, Sign Codes, and the Consumption Law
view article Subject Bucketing: Teaching a Diffusion Model New Prompt Languages Without Forgetting
AbstractPhil
• • 1
view article geolip-aleph-void: The First Relational Geometric Vocabulary Patchwork
view article Reading the Voids: Topological Contribution Signals in Frozen Geometric Codebooks
view article Fused Batched Thin SVD, Part II: Extending the Jacobi Pipeline to N=6 with Configurable Convergence
view article H2 Omega Confirmed, Paradigm Shift: Attempting to Disprove Omega As A Whole
AbstractPhil
• • 1
view article The Polygonal Omega: Trained Sphere-Solvers Are Projective Codebooks
AbstractPhil
• • 1
view article Three Geometric Bands in a Sphere-Normalized Patch Autoencoder
view article The Geometric Engine: Structural Attractors in Neural Network Weight Space
view article FL Hybrid Eigendecomposition Beating cuSOLVER's Mathematical Purity with Compilable PyTorch
view article Ryan Spearman: Geometric Variant Effect Prediction Through Quaternion-Composed Dual Expert Alignment
view article Fused Batched Thin SVD: Engineering a 5000× Speedup with Triton Kernels