FrontierLab-Governance(external site)
Shared rulebook and gate software
- 2 open PRs
- 5 merged PRs
- 30 commits
- protocol 2.0.0
- CI green
Work · Research governance
"This repository manages the research process. It does not decide what is true about nature or mathematics."
Frontier Lab takes the same method to open science. No model is exempt from checks because of its brand, and two models agreeing is not independent verification. The tooling is the Python standard library only, and each round's default budget is $0. Status is shown as it is: most problems are still OPEN; the Physics, Biology and Chemistry results are still in review; Math has one merged round of evidence and claims no resolution.
It is not a benchmark or a Q&A set. It is a multi-model research workflow under strict governance.
The whole program was set up between 2026-09-27 and 2026-09-29.
Governance on the left; one row per lab, one cell per problem card. ◆ marks each lab's first-round problem.
| 001 | 002 | 003 | 004 | 005 | 006 | 007 | 008 | 009 | 010 | ||
|---|---|---|---|---|---|---|---|---|---|---|---|
| FrontierLab-Governance | FrontierMath | First round: MATH-001 | MATH-002 | MATH-003 | MATH-004 | MATH-005 | MATH-006 | MATH-007 | MATH-008 | MATH-009 | MATH-010 |
| FrontierPhysics | First round: PHYS-001 | PHYS-002 | PHYS-003 | PHYS-004 | PHYS-005 | PHYS-006 | PHYS-007 | PHYS-008 | PHYS-009 | PHYS-010 | |
| FrontierBiology | First round: BIO-001 | BIO-002 | BIO-003 | BIO-004 | BIO-005 | BIO-006 | BIO-007 | BIO-008 | BIO-009 | BIO-010 | |
| FrontierChemistry | CHEM-001 | CHEM-002 | CHEM-003 | First round: CHEM-004 | CHEM-005 | CHEM-006 | CHEM-007 | CHEM-008 | CHEM-009 | CHEM-010 | |
| FrontierComputerScience | First round: CS-001 | CS-002 | CS-003 | CS-004 | CS-005 | CS-006 | CS-007 | CS-008 | CS-009 | CS-010 | |
| FrontierStatistics | First round: STAT-001 | STAT-002 | STAT-003 | STAT-004 | STAT-005 | STAT-006 | STAT-007 | STAT-008 | STAT-009 | STAT-010 | |
| FrontierMetaScience | First round: META-001 | META-002 | META-003 | META-004 | META-005 | META-006 | META-007 | META-008 | META-009 | META-010 | |
| FrontierSocialScience | SOC-001 | SOC-002 | SOC-003 | SOC-004 | SOC-005 | SOC-006 | SOC-007 | First round: SOC-008 | SOC-009 | SOC-010 | |
| FrontierMaterials | MAT-001 | MAT-002 | First round: MAT-003 | MAT-004 | MAT-005 | MAT-006 | MAT-007 | MAT-008 | MAT-009 | MAT-010 | |
| FrontierAstronomy | ASTRO-001 | ASTRO-002 | First round: ASTRO-003 | ASTRO-004 | ASTRO-005 | ASTRO-006 | ASTRO-007 | ASTRO-008 | ASTRO-009 | ASTRO-010 | |
| FrontierEarth | EARTH-001 | EARTH-002 | First round: EARTH-003 | EARTH-004 | EARTH-005 | EARTH-006 | EARTH-007 | EARTH-008 | EARTH-009 | EARTH-010 | |
| FrontierNeuroscience | First round: NEURO-001 | NEURO-002 | NEURO-003 | NEURO-004 | NEURO-005 | NEURO-006 | NEURO-007 | NEURO-008 | NEURO-009 | NEURO-010 | |
| FrontierEconomics | First round: ECON-001 | ECON-002 | ECON-003 | ECON-004 | ECON-005 | ECON-006 | ECON-007 | ECON-008 | ECON-009 | ECON-010 | |
| FrontierEngineering | ENG-001 | ENG-002 | ENG-003 | First round: ENG-004 | ENG-005 | ENG-006 | ENG-007 | ENG-008 | ENG-009 | ENG-010 | |
| FrontierMedicine | First round: MED-001 | MED-002 | MED-003 | MED-004 | MED-005 | MED-006 | MED-007 | MED-008 | MED-009 | MED-010 |
a fresh four-way search within 24 hours of starting: general, discipline, solution, criticism. If an outside solution exists, it is recorded as COMPLETED_EXTERNAL and the work stops.
reproduce the known result first.
the problem and its verifier are frozen first.
default budget per round: 30 minutes, $0, 100 tries.
no author signs off their own independent verification.
results and failures both stay; records can be added to, never deleted.
authors never merge their own work; the owner decides.
"No agent is exempt from verification because of its brand."
"Two models agreeing is not an independent experiment."
OPEN means only this: this bounded search found no confirmed solution of the same scope. It does not mean "unsolved".
OPENPARTIALCLAIMED_RESOLVEDCOMPLETED_EXTERNALCOMPLETED_INTERNALPAUSEDRETRACTEDDRAFTADMITTEDPAUSEDCLOSED_EXTERNALFINISHEDShared rulebook and gate software
Exact constructions → certificates → Lean proofs
FrontierMath is not affiliated with Epoch AI's FrontierMath benchmark.
Reproducible baseline → physical consistency → predictions that tell models apart
Public data → baseline → cross-condition validation → testable hypotheses
Keeps computed and experimental references separate; starts from solvation
Formal verification and honest evaluation of AI coding agents
Valid inference under distribution shift, selection and dependence
Tests whether AI-agent science and its governance actually improve research
Measurement → identification → replication → transportability
Separates what computation predicts from what can actually be synthesized
Treats catalogs and their selection functions as research objects
Calibrating forecasts of rare, extreme events
Designs stimuli that make brain-computation models disagree
Measures what AI actually does to firms, tasks and work
Simulation-first: fault injection and uncertainty before any deployment
Validating medical AI across hospitals and over time
All 431,008 configurations in the public Flammenkamp database for n = 2..76 were verified legal; n = 75 is the only gap.
problems/MATH-001/experiments/canonical_evidence/CANONICAL_FACTS.mdAs of Records for n = 71–74 and 76 were cross-checked through two independent retrieval paths; an independent third verifier was checked by a 429-case adversarial suite.
Each column is one n (2 to 76); the empty one is n = 75. Column height carries no data.
Lean 4 CI is pinned to v4.34.1. The only theorem so far is a toolchain smoke test, "not a frontier result".
FrontierMath is not affiliated with Epoch AI's FrontierMath benchmark.
These three results sit in PRs that haven't been merged.
Physics PR #4 reproduces another model's structure-function estimator exactly (a 1.1e-14 gap), records a frozen E3 FAIL, and retracts an "asymptotic saturation" claim. It makes no claim about real Navier–Stokes turbulence.
Biology PR #2 withdraws an unverified PASS. It found 3,929 train/test overlaps, re-estimated with a donor split at 0.1616, and kept the historical failures.
Chemistry PR #1 adds a FreeSolv validator and reproduces the GAFF anchor (1.114); the frozen C4 and C5 thresholds are recorded as FAIL. PR #2 fixes a CRLF/LF hash mismatch and marks C4 as post-hoc.
FrontierChemistry/pull/1(external site)FrontierChemistry/pull/2(external site)
frontier.py 796 lines, standard library onlyCommands
validatestartadmitcheck-diffcheck-pinsdecideFrontierLab-Governance/tools/frontier.py(external site)FrontierLab-Governance/STATUS.md(external site)
Protocol 1.0.0 → 1.1.0 → 1.2.0 → 2.0.0 (breaking). The 11 newer labs pin 2.0.0; the first 4 still pin 1.0.0.
Roles include Explorer, Verifier, Skeptic and Integrator. No author signs off their own independent verification; the owner decides merges.
All 11 new labs failed CI on their first run; the tooling was fixed and re-pinned. "The failure was not deleted or rewritten as a first-time pass."
FrontierLab-Governance/EXPANSION-2026-09-28.md(external site)
Branch protection is only a proposal; it isn't enabled.
FrontierLab-Governance/BRANCH_PROTECTION_PROPOSAL.md(external site)
11 of the 15 labs have no research rounds yet.
Branch protection is only a proposal; it isn't enabled.
FrontierMath's STATUS.md is stale and still says 0 rounds.
None of the 16 repos has a live demo yet (Pages returns 404).