Li Jiabao
Static · 0 glimmers中文 以中文閱讀本頁

Work · Research governance

Frontier Lab

"This repository manages the research process. It does not decide what is true about nature or mathematics."

Frontier Lab takes the same method to open science. No model is exempt from checks because of its brand, and two models agreeing is not independent verification. The tooling is the Python standard library only, and each round's default budget is $0. Status is shown as it is: most problems are still OPEN; the Physics, Biology and Chemistry results are still in review; Math has one merged round of evidence and claims no resolution.

It is not a benchmark or a Q&A set. It is a multi-model research workflow under strict governance.

  • 1 governance repo
  • 15 labs
  • 150 problem cards
  • Python standard library
  • $0 default per round

The whole program was set up between 2026-09-27 and 2026-09-29.

Real records from the 2026-09-30 snapshot. One point per record.1 ○ = 1 problem card

Governance on the left; one row per lab, one cell per problem card. ◆ marks each lab's first-round problem.

10 problem cards per lab
001002003004005006007008009010
FrontierLab-GovernanceFrontierMathFirst round: MATH-001MATH-002MATH-003MATH-004MATH-005MATH-006MATH-007MATH-008MATH-009MATH-010
FrontierPhysicsFirst round: PHYS-001PHYS-002PHYS-003PHYS-004PHYS-005PHYS-006PHYS-007PHYS-008PHYS-009PHYS-010
FrontierBiologyFirst round: BIO-001BIO-002BIO-003BIO-004BIO-005BIO-006BIO-007BIO-008BIO-009BIO-010
FrontierChemistryCHEM-001CHEM-002CHEM-003First round: CHEM-004CHEM-005CHEM-006CHEM-007CHEM-008CHEM-009CHEM-010
FrontierComputerScienceFirst round: CS-001CS-002CS-003CS-004CS-005CS-006CS-007CS-008CS-009CS-010
FrontierStatisticsFirst round: STAT-001STAT-002STAT-003STAT-004STAT-005STAT-006STAT-007STAT-008STAT-009STAT-010
FrontierMetaScienceFirst round: META-001META-002META-003META-004META-005META-006META-007META-008META-009META-010
FrontierSocialScienceSOC-001SOC-002SOC-003SOC-004SOC-005SOC-006SOC-007First round: SOC-008SOC-009SOC-010
FrontierMaterialsMAT-001MAT-002First round: MAT-003MAT-004MAT-005MAT-006MAT-007MAT-008MAT-009MAT-010
FrontierAstronomyASTRO-001ASTRO-002First round: ASTRO-003ASTRO-004ASTRO-005ASTRO-006ASTRO-007ASTRO-008ASTRO-009ASTRO-010
FrontierEarthEARTH-001EARTH-002First round: EARTH-003EARTH-004EARTH-005EARTH-006EARTH-007EARTH-008EARTH-009EARTH-010
FrontierNeuroscienceFirst round: NEURO-001NEURO-002NEURO-003NEURO-004NEURO-005NEURO-006NEURO-007NEURO-008NEURO-009NEURO-010
FrontierEconomicsFirst round: ECON-001ECON-002ECON-003ECON-004ECON-005ECON-006ECON-007ECON-008ECON-009ECON-010
FrontierEngineeringENG-001ENG-002ENG-003First round: ENG-004ENG-005ENG-006ENG-007ENG-008ENG-009ENG-010
FrontierMedicineFirst round: MED-001MED-002MED-003MED-004MED-005MED-006MED-007MED-008MED-009MED-010

One round, seven gates

  1. Check whether it's solved

    a fresh four-way search within 24 hours of starting: general, discipline, solution, criticism. If an outside solution exists, it is recorded as COMPLETED_EXTERNAL and the work stops.

  2. Reproduce

    reproduce the known result first.

  3. Freeze

    the problem and its verifier are frozen first.

  4. Bounded exploration

    default budget per round: 30 minutes, $0, 100 tries.

  5. Independent verification

    no author signs off their own independent verification.

  6. Record

    results and failures both stay; records can be added to, never deleted.

  7. Reviewed merge

    authors never merge their own work; the owner decides.

Failures on record

Rules

  1. "No agent is exempt from verification because of its brand."

  2. "Two models agreeing is not an independent experiment."

OPEN

What OPEN means

OPEN means only this: this bounded search found no confirmed solution of the same scope. It does not mean "unsolved".

Problem states
  • OPEN
  • PARTIAL
  • CLAIMED_RESOLVED
  • COMPLETED_EXTERNAL
  • COMPLETED_INTERNAL
  • PAUSED
  • RETRACTED
Round states
  • DRAFT
  • ADMITTED
  • PAUSED
  • CLOSED_EXTERNAL
  • FINISHED

15 labs, one governance core

  • Merged1
  • In review3
  • No research rounds yet11

10 problem cards per labData as of 2026-09-30

View
  • Merged

    FrontierMath(external site)

    Exact constructions → certificates → Lean proofs

    • 1 merged round
    • Not claimed
    • 6 open PRs
    • 2 merged PRs
    • 8 commits
    • protocol 1.0.0
    • CI green
    • First round: MATH-001

    FrontierMath is not affiliated with Epoch AI's FrontierMath benchmark.

  • In review

    FrontierPhysics(external site)

    Reproducible baseline → physical consistency → predictions that tell models apart

    • 5 open PRs
    • 2 commits
    • protocol 1.0.0
    • CI green
    • First round: PHYS-001
  • In review

    FrontierBiology(external site)

    Public data → baseline → cross-condition validation → testable hypotheses

    • 2 open PRs
    • 2 commits
    • protocol 1.0.0
    • CI green
    • First round: BIO-001
  • In review

    FrontierChemistry(external site)

    Keeps computed and experimental references separate; starts from solvation

    • 2 open PRs
    • 2 commits
    • protocol 1.0.0
    • CI green
    • First round: CHEM-004
  • No research rounds yet

    FrontierStatistics(external site)

    Valid inference under distribution shift, selection and dependence

    • 2 merged PRs
    • 7 commits
    • protocol 2.0.0
    • CI green
    • First round: STAT-001
  • No research rounds yet

    FrontierMetaScience(external site)

    Tests whether AI-agent science and its governance actually improve research

    • 2 merged PRs
    • 7 commits
    • protocol 2.0.0
    • CI green
    • First round: META-001
  • No research rounds yet

    FrontierSocialScience(external site)

    Measurement → identification → replication → transportability

    • 2 merged PRs
    • 7 commits
    • protocol 2.0.0
    • CI green
    • 10 OPEN
    • First round: SOC-008
  • No research rounds yet

    FrontierMaterials(external site)

    Separates what computation predicts from what can actually be synthesized

    • 2 merged PRs
    • 7 commits
    • protocol 2.0.0
    • CI green
    • 10 OPEN
    • First round: MAT-003
  • No research rounds yet

    FrontierAstronomy(external site)

    Treats catalogs and their selection functions as research objects

    • 2 merged PRs
    • 7 commits
    • protocol 2.0.0
    • CI green
    • 10 OPEN
    • First round: ASTRO-003
  • No research rounds yet

    FrontierEarth(external site)

    Calibrating forecasts of rare, extreme events

    • 2 merged PRs
    • 7 commits
    • protocol 2.0.0
    • CI green
    • 10 OPEN
    • First round: EARTH-003
  • No research rounds yet

    FrontierNeuroscience(external site)

    Designs stimuli that make brain-computation models disagree

    • 2 merged PRs
    • 7 commits
    • protocol 2.0.0
    • CI green
    • 10 OPEN
    • First round: NEURO-001
  • No research rounds yet

    FrontierEconomics(external site)

    Measures what AI actually does to firms, tasks and work

    • 2 merged PRs
    • 7 commits
    • protocol 2.0.0
    • CI green
    • 10 OPEN
    • First round: ECON-001
  • No research rounds yet

    FrontierEngineering(external site)

    Simulation-first: fault injection and uncertainty before any deployment

    • 2 merged PRs
    • 7 commits
    • protocol 2.0.0
    • CI green
    • 10 OPEN
    • First round: ENG-004
  • No research rounds yet

    FrontierMedicine(external site)

    Validating medical AI across hospitals and over time

    • 2 merged PRs
    • 7 commits
    • protocol 2.0.0
    • CI green
    • 10 OPEN
    • First round: MED-001

Focus: no-three-in-line (MATH-001)

All 431,008 configurations in the public Flammenkamp database for n = 2..76 were verified legal; n = 75 is the only gap.

431,008Source: FrontierMath problems/MATH-001/experiments/canonical_evidence/CANONICAL_FACTS.mdAs of

Records for n = 71–74 and 76 were cross-checked through two independent retrieval paths; an independent third verifier was checked by a 429-case adversarial suite.

A diagram, not data.

Each column is one n (2 to 76); the empty one is n = 75. Column height carries no data.

n = 75 · gap

Not claimed

  • No claim that D(75) = 150.
  • No UNSAT certificate.
  • MATH-001 is not resolved.

Lean 4 CI is pinned to v4.34.1. The only theorem so far is a toolchain smoke test, "not a frontier result".

FrontierMath is not affiliated with Epoch AI's FrontierMath benchmark.

In review

These three results sit in PRs that haven't been merged.

  • In review

    Physics PR #4 reproduces another model's structure-function estimator exactly (a 1.1e-14 gap), records a frozen E3 FAIL, and retracts an "asymptotic saturation" claim. It makes no claim about real Navier–Stokes turbulence.

  • In review

    Biology PR #2 withdraws an unverified PASS. It found 3,929 train/test overlaps, re-estimated with a donor split at 0.1616, and kept the historical failures.

  • In review

    Chemistry PR #1 adds a FreeSolv validator and reproduces the GAFF anchor (1.114); the frozen C4 and C5 thresholds are recorded as FAIL. PR #2 fixes a CRLF/LF hash mismatch and marks C4 as post-hoc.

The governance core

  • frontier.py 796 lines, standard library only
  • 168 tests
  • mutation 29/29
  • 30 commits
  • 7 PRs (2 open)

Commands

  • validate
  • start
  • admit
  • check-diff
  • check-pins
  • decide

FrontierLab-Governance/tools/frontier.py(external site)FrontierLab-Governance/STATUS.md(external site)

  1. 1.0.0
    • MATH
    • PHYS
    • BIO
    • CHEM
  2. 1.1.0
  3. 1.2.0
  4. 2.0.0
    • CS
    • STAT
    • META
    • SOC
    • MAT
    • ASTRO
    • EARTH
    • NEURO
    • ECON
    • ENG
    • MED

Protocol 1.0.0 → 1.1.0 → 1.2.0 → 2.0.0 (breaking). The 11 newer labs pin 2.0.0; the first 4 still pin 1.0.0.

Roles include Explorer, Verifier, Skeptic and Integrator. No author signs off their own independent verification; the owner decides merges.

All 11 new labs failed CI on their first run; the tooling was fixed and re-pinned. "The failure was not deleted or rewritten as a first-time pass."

FrontierLab-Governance/EXPANSION-2026-09-28.md(external site)

Branch protection is only a proposal; it isn't enabled.

FrontierLab-Governance/BRANCH_PROTECTION_PROPOSAL.md(external site)

Each lab's safety limits

Biology
benign public benchmarks only; no pathogens, toxins or wet-lab work; no clinical diagnosis or treatment advice.
Chemistry
refuses weapons, toxic agents, explosives and wet-lab execution.
Medicine
no diagnostic or treatment conclusions for individuals.
Engineering
no connection to real grids, traffic, robots or industrial control.
Earth
public data and backtests only; no real-time hazard guidance for individuals.
Neuroscience
no invasive work, no personal neural data.
Materials
paid DFT/GPU runs need separate authorization; the default budget is $0.

What didn't get done

  1. 11 of the 15 labs have no research rounds yet.

  2. Branch protection is only a proposal; it isn't enabled.

  3. FrontierMath's STATUS.md is stale and still says 0 rounds.

  4. None of the 16 repos has a live demo yet (Pages returns 404).