Anonymous research results · ledger as of 2026-08-02

One operator.
Measured exactly.

GEML asks whether a graph neural network can learn symbolic mathematics from a representation in which every internal operation is the same gate. We report the complete authenticated 250,000-row corpus and the bounded Phase-B evidence, including nulls and unfinished gates.

This is a results page, not a completion claim. Goals 1–3 are authenticated at full corpus scale; Goals 4–5 are authenticated at their explicitly reported study denominators. Phase B contains a three-cell Goal 6 source refresh, an incomplete Goal 7 grid, a retained null Goal 8 value head, and no authenticated production evidence for Goals 9, 11, or 12.

250,000unique corpus rows, split 175k / 25k / 25k / 25k
40.6602×median raw pure-EML expansion α
952.1371×mean raw pure-EML expansion α
39.375×mean recovery from exact EML-DAG sharing
44.5127%frequent-motif IID MDL savings; reconstruction failures: 0
Limited Phase-B reading: two of three current-source Goal 6 refresh seeds fall below the pilot-derived 0.577 constant-majority BCE reference (losses 2.391449, 0.526723, 0.538694). This is a three-cell source refresh, not the six-arm Goal 6 verdict.
01 / METHOD

A controlled representation study

Source expressions are compiled with the authoritative official_v4 formulas into pure EML trees. Exact sharing and verifier-gated transformations create representation channels; bounded learning tests then measure what topology buys and what it costs. The representation is intentionally austere so that a result can be attributed to structure rather than a growing operator vocabulary.

SourceAST expressions with domain and family metadata
Compileofficial-v4 pure EML tree, canonical and auditable
Shareexact DAG construction by hash-consing
Compresse-graph, frequent motif, and learned motif channels
Testbounded GNN, rewrite, proof, and regression tracks

The AST-DAG channel is the fairness baseline. It separates the effect of graph sharing from the effect of the EML abstraction itself.

02 / CORPUS

Goals 1–5: corpus, compression, and exact costs

Goal 1 · source corpus complete

Exactly 250,000 unique source expressions were accepted from 286,413 attempts with no unsupported rows and no cross-split identity collisions.

  • Splits: 175,000 / 25,000 / 25,000 / 25,000 train, validation, test-IID, test-OOD.
  • Six families and 1–6 variables; trig and hyperbolic operators are present under their approved domains.
  • Authoritative s-expressions, expression IDs, and structural identities are each unique at 250,000.

Goal 2 · official-v4 expansion complete

Official-v4 compilation succeeded for all 250,000 / 250,000 rows. The semantic audit is a selected subset, not a conversion-failure count.

  • Selected 280; materialized 273: 203 passed, 3 mismatched, 45 nonfinite, 22 overflow, and 7 exceeded the node limit before materialization.
  • Raw α: median 40.6602, mean 952.1371, approximate p99 10,448.6.
  • Zero rows fall below the preregistered 1.29–1.50 threshold range.

Goal 3 · exact DAGs complete

Both AST-DAG and EML-DAG were built for all 250,000 / 250,000 rows. Exact EML sharing recovers 39.375× of the raw EML tree on mean node count.

  • Mean EML-DAG / AST-tree ratio: 8.3344; EML-DAG / AST-DAG: 10.4750.
  • No row is structurally competitive with the AST baseline; the best remaining ratio is 8/7.

Goal 4 · verifier-gated e-graphs complete

Each mode processed 30,000 rows. Exactly 18,210 were costed; 11,790 were retained as unsupported or independently invalid, with no timeouts.

  • safe_real: 4,349 improvements, 23.8825% of costed rows.
  • positive_real_formal: 5,026 improvements, 27.6002% of costed rows.
  • Failures are principally unsupported trig/hyperbolic operators or validation rejection, not hidden zeros.

Goal 5 · motifs and learned baselines complete

The frequent motif dictionary is the strongest simple lossless channel: 44.5127% IID and 43.6983% OOD MDL savings with zero reconstruction failures.

  • Learned motifs lose to equal-budget frequent motifs (IID: 324.5M vs 317.7M bits).
  • The neural ranker is faster, but loses exact-best selection: 8,349/10,752 IID groups versus 8,796 for the EML-tree heuristic and 8,661 for AST-DAG.
  • Production export covers 250,000 expressions, 1,250,000 graph views, 250,000 hierarchies, and 2,500 batches with zero validation or reconstruction failures.
  • These are retained null comparisons, not silently omitted wins.
Family-level raw EML expansion chart showing median alpha versus preregistered thresholds across six corpus families.
Figure 1. Expansion is family-dependent and heavy-tailed. The corpus-wide median is 40.6602× while the mean is 952.1371×; neither is substituted for the other.
Two-panel chart: independent structural compression ratios for raw and DAG representations, and Goal 4 outcome composition for safe-real and positive-real-formal modes.
Figure 2. Exact sharing recovers 39.375× of raw EML tree size, but EML-DAG remains larger than the AST baselines. The outcome panel reports Goal 4 row composition; e-graph improvements use the costed denominator.
Learned motif and neural ranker baseline comparison chart with frequent motifs and structural heuristics.
Figure 3. Frequent motifs remain the practical lossless baseline. Learned selection and neural ranking are useful comparisons, including their retained negative results.
03 / PHASE B

Authenticated GPU evidence, with boundaries

The archived Phase-B package is operationally authenticated, but operational completion is not the same as a completed scientific gate. The table below reports only the cells and metrics present in the ledger.

Goal 6 · equivalence source refresh

partial · source refresh only Three current-source pure-EML cells were run with seeds 20260726–20260728. They do not constitute the planned six-arm, three-seed (18-cell) Goal 6 verdict; the measured result is a pure-EML viability signal, not a general learning verdict.

Reviewer shorthand: pure-EML viability signal — 2/3 source-refresh seeds below the pilot majority floor; controlled six-arm comparison unrun.

SeedValidation BCEWall timeBelow pilot 0.577 reference?
202607262.3914492883.6 sno
202607270.5267232890.9 syes
202607280.5386942888.9 syes

Goal 7 · rewrite-step prediction

partial · invalids retained The logical denominator is 18 cells: 13 complete + 5 invalid. The invalid cells remain in the denominator. A separate retrieval auxiliary grid completed 15 / 15 operational cells; this is operational completeness only, not a retrieval-quality verdict and not a substitute for the incomplete scheduler.

Goal 8 · value-head diagnostics

retained null / collapsed Three runs completed with 1,500 optimizer steps and 452,820 trainable parameters. Validation MAE ranges from 0.8296–0.8472, but held-out ranking is at chance and OOD Spearman is undefined; the value head must not be presented as a useful ranker.

SeedValidation MAEIID SpearmanOOD Spearman
202607260.8472−0.0142undefined
202607270.8376+0.0021undefined
202607280.8296−0.0089undefined
Equivalence viability chart showing validation losses for three source-refresh seeds and the pilot-derived constant-majority reference.
Figure 4. The three-cell source refresh is mixed: two seeds fall below the pilot-derived 0.577 constant-majority reference and one does not. The missing arms are not imputed.
Value-head diagnostics showing validation MAE and near-zero IID Spearman correlations for three seeds.
Figure 5. Value-head error alone looks plausible; rank correlation does not. Near-zero IID Spearman and undefined OOD Spearman support the retained null/collapsed interpretation.
04 / LEDGER

Goal 1–12 status matrix

Statuses describe authenticated evidence, not optimism about an implementation path. “Unavailable / unrun” is distinct from zero, and “partial” is distinct from complete.

completepartialretained nullbounded gate failunavailable / unrun
GoalStatusEvidence-safe interpretation
1complete250,000 unique source rows; all planned splits and corpus QA are authenticated.
2complete250,000 / 250,000 official-v4 compilations; selected semantic audit caveats retained.
3completeExact AST-DAG and EML-DAG corpus build; no structurally competitive EML row.
4completeTwo 30k mode runs; improvements reported over costed denominators.
5completeFrequent motif wins simple lossless comparison; learned channels retain null results.
6partialThree-cell current-source refresh only; not the six-arm Goal 6 verdict.
7partial13 complete + 5 invalid of 18 logical cells; retrieval 15/15 is separate.
8retained nullThree completed value runs; collapsed ranking signal and undefined OOD Spearman.
9unavailable / unrunNo authenticated production symbolic-regression comparison.
10bounded gate fail74-row CPU compiler gate retained; eight asin/acos endpoint cells fail the bounded endpoint evaluation. Broader production was not authenticated.
11unavailable / unrunNo authenticated scale-up or external-LLM benchmark ledger.
12unavailable / unrunNo authenticated final-consolidation or release gate.
05 / REPRODUCE

Evidence, provenance, and limits

Every headline number on this page is copied from the compact, machine-readable results ledgers or the linked goal summaries. The JSON copies are page-local so the public page remains self-contained when deployed from the docs/ source directory.

Machine-readable ledgersresults.json and detailed_results.json preserve fields, denominators, statuses, and warnings.
Human-readable provenancePROVENANCE.md, DETAILED_RESULTS.md, and the goal summaries explain source and interpretation.
What is not shippedLarge raw corpora, checkpoints, GPU logs, and archives are excluded from this anonymous page release. Their absence is not a claim that the associated gate is complete.

No missing value is imputed. Checksums and artifact authentication are documented in the linked provenance and validation records. GPU bitwise reproducibility across hardware is not claimed; retained cuBLAS warnings remain part of the evidence record.

Limitations and non-claims