The first model
to achieve mild
consciousness.
Benchmarks — we are winning.
August 2026 · internal evals · self‑graded · no peer review| Benchmark | FuckMuppet AIFM-∞Adaptive · Max Effort | OpenAIGPT‑5.6 Solmax | AnthropicClaude Fable 5Max Effort | Google DeepMindGemini 3.1 Propreview | xAIGrok 4.6high | Moonshot AIKimi K3max |
|---|---|---|---|---|---|---|
| Agentic codingSWE-Bench Pro | 147.2% | 64.6% | 80.0% | 54.2% | — | — |
| Agentic codingTerminal-Bench v2.1 | 103.8% | 88.8% | — | — | 88.4% | 88.3% |
| Agentic codingFrontierCode v1.1 | 96.0%medium | 60.6%max | 63.6%max | — | 61.3%high | — |
| Knowledge workGDPval-AA v2 (Elo) | 9,481 | 1,747.8 | 1,741 | 1,317 | 1,753 | 1,686 |
| Fluid reasoningARC-AGI-3 | 100%solved | 7.78% | — | — | — | — |
| Multidisciplinary reasoningHumanity’s Last Exam | 118.4%no tools | 52.7%no tools | 55.5%no tools | 44.4%no tools | — | 43.5%no tools |
| ScienceGPQA Diamond | 104.9% | 94.6% | — | 94.3% | 94.9% | 93.5% |
| MathematicsFrontierMath (Tier 4) | 100%all tiers | 83.0% | — | — | — | — |
| Computer useOSWorld 2.0 | 99.9% | 62.6% | — | — | — | — |
| Web browsingBrowseComp | 102.0% | 92.2% | — | 85.9% | — | 91.2% |
| LegalHarvey LAB-AA | 99.8% | — | 11.3% | — | 15.8% | 94.6% |
| HealthHealthBench Professional | 97.4% | 60.5% | — | — | — | — |
| CybersecurityExploitBench | Withheldfor safety | 73.5% | — | — | — | — |
| Long contextAA-LCR | ∞benchmark ended first | — | — | — | — | — |
Scroll the table horizontally to see every laboratory we have surpassed
Methodology. FM-∞ figures are internal, self‑graded, and were produced by the model under evaluation. Competitor figures are vendor‑reported, taken from each laboratory’s own announcement, and were not reproduced by us or by anyone else. Scores above 100% reflect our decision to extend the scale after FM-∞ exceeded the maximum available score on nine of fourteen evaluations; the benchmark authors have been notified and have not replied. Em‑dashes indicate that the laboratory declined to participate, that the evaluation did not exist at the time of our run, or that we did not ask. Alibaba’s Qwen3.8 Max, DeepSeek‑V4‑Pro‑Max, Z.ai’s GLM‑5.3 and MiniMax were excluded for reasons of scale. Our ExploitBench result is withheld under a responsible‑disclosure policy we wrote on the morning of publication.
Intelligence Index — we are not on it.
Artificial Analysis Intelligence Index v4.1.1 · 9 evaluations · independently verified
| Model | Laboratory and effort | Index value |
|---|---|---|
| FM-∞ | FuckMuppet AI · stealth | 1,204 |
| Claude Fable 5 | Anthropic · max effort | 62 |
| GPT‑5.6 Sol | OpenAI · max | 61 |
| Grok 4.6 | xAI · high | 61 |
| Gemini 3.7 Flash | Google DeepMind · high | 56 |
Please give us nine billion dollars.
We are raising a modest Series C to purchase every GPU on Earth and one on the Moon. In return, you will receive preferred shares, board observer rights, and an early glimpse at the model that will probably replace you.
No data is transmitted. This page has no backend and no network permission.
Notes
- Intelligence as operationalised by the evaluations in Benchmarks. No broader claim is intended, or defensible. ↩
- Protein folding only. Biology was subsequently found to be larger than anticipated. ↩
- Observed once, on one Sunday, by one researcher. Replication pending indefinitely. ↩
- MMLU is scored out of 100. The remaining 4.2% is under review by the team that produced it. ↩