Jump to content

Draft:Millennium Research

From Wikipedia, the free encyclopedia

Millennium Research Inc. is an American artificial intelligence company that develops human-verified faithfulness evaluations and certified datasets for formal mathematics. The company builds tooling that measures whether formal statements written in the Lean 4 proof assistant faithfully capture the meaning of the informal mathematics they represent — a property that neither the Lean compiler nor, in the company's published audits, frontier large language model judges reliably determine on their own.[1]

History

[edit]

Millennium Research was founded by Shayaan Siddique and Ibrahim, who serve as co-founders. Siddique, the company's chief technology officer, leads technical architecture; Ibrahim leads research. The company was incorporated as a Delaware C corporation and raised pre-seed funding from venture firm 3kVC.[citation needed]

Products and research

[edit]

Faithfulness evaluation suite

[edit]

The company's core product is an evaluation system for detecting faithfulness defects — cases where a formal Lean 4 statement compiles and type-checks but does not correspond to the source mathematics. The system combines two frontier LLM judges under strict consensus, a numeric counterexample probe, and a detection ladder attributing each caught defect to the layer that found it. Judge error rates are calibrated against a frozen set of 886 human verdicts and published with confidence intervals.[1]

As of 2026, the company reports having audited 3,299 formal statements across three public benchmarks and four open autoformalization systems as a full census, identifying 156 human-confirmed faithfulness defects with a pooled screen precision of 91.8% under human review. It reports that all 139 counterexample-backed flags to date have been upheld on human review, and 78.3% recall measured against the independent miniF2F v2 correction effort. Defect filings for the miniF2F, miniF2F v2, and ProofNet# benchmarks are public.[1]

Certified dataset

[edit]

The company produces a dataset of informal–formal statement pairs from open-licensed mathematics (CC0 and CC-BY), machine-verified in Lean 4 against mathlib and certified faithful by human reviewers. A core design invariant is that automated screens may only reject candidate pairs; certification is reserved for humans. The certified product excludes proof text, is license-gated at ingest, and maintains a tamper-evident history.[1]

Positioning

[edit]

Millennium Research positions its work as evaluation and reward infrastructure for laboratories training theorem-proving models with reinforcement learning, where unfaithful formal statements can act as a vector for specification gaming. The company states a long-term goal of contributing to machine-checked progress on open mathematical problems.[1]

See also

[edit]

References

[edit]
  1. 1 2 3 4 5 "Millennium Research". millenniumresearch.ai. Retrieved 21 July 2026.
[edit]