


Astrophysics PhD work at Harvard. Directed UC Berkeley's Center for Cosmological Physics. Head of Data Science at Google, Senior Data Scientist at SpaceX. Founder & CEO of Climax Foods.
Surgeon and neuroscientist. Director, Frontier Technologies Lab, Cambridge.
Engineering Lead. Machine learning and materials science, UC Berkeley.
Lab Integration Lead. Molecular engineering, UChicago; founded a DOE-funded energy startup.
Google DeepMind. Speak with them under NDA; joins full-time on close.
Google DeepMind. Speak with them under NDA; joins full-time on close.
OpenAI. Speak with them under NDA; joins full-time on close.
Investors can speak with the incoming frontier-lab researchers under NDA today; they join full-time when the round closes. Six papers, five on arXiv; seven patent families, with the continual-learning mechanisms in filing now.
What sets us apart is the evidence behind the claims: matched protocols, multiple seeds, published kill conditions, and nulls reported on the slide. Diligence gets a reproduction package.
ex-Director of Research, Google; co-author, AI: A Modern Approach
Neuroscientist, Stanford.
Science of science, UChicago.
Cosmologist, Institute for Advanced Study.
Sherpa Capital; early Uber, Airbnb.
Google DeepMind.
CIO, SWIFT.
Former Head of Growth, TikTok.
Former CEO, Tyson Foods.
The wedge is research organizations, entered from the bottom up: researchers adopt the system inside the AI sessions where their work already happens. Free client, Pro and Team seats, Enterprise licenses; revenue starts as subscription, royalties on transacted value later. Signed LOIs with two major scientific publishers sit outside this plan while those partners are paused. The full pricing table is in the appendix.
The entire system can run on customer hardware with no data leaving the building; for regulated industry and enterprise outside the United States that is often the deciding requirement, and no closed-weights lab can offer it.
The ask. A seed round weighted to senior research and engineering hires, funding three milestones: the first paid cross-organization deployments, the weight path productized at the scale this deck states, and the network reaching self-sustaining density. Sapience is at its most useful on day one, alone; the network is not the price of admission but the compounding on top.
The core plan is people, on roughly 24 months of runway; we serve open weights rather than pretraining, so compute is ~12% of the raise rather than ~70%. Sizing detail in the appendix.
"In today's systems, we train them, and then they're kind of frozen and put out into the world. What you'd like is for those systems to continually learn online from experience."Demis Hassabis · India AI Impact Summit · 2026
Sapience resolves both. It learns continuously with no retraining, and attribution at write time lets scientists contribute while keeping ownership and credit. One mechanism serves scientists and builds the pre-publication dataset frontier AI currently lacks.
The knowledge object itself carries what the mechanisms need: provenance anchors it to reality, typed links carry how it connects, and the record stays current as findings replace one another. The nine mechanisms operate on the object itself (there is no separate database underneath).

The visible banding on the right is the constraint made real. A low-rank update repeats a few patterns across the whole matrix instead of rewriting it freely, so it moves the model along a handful of directions and stays removable. The store decides what flows through that narrow channel each night. The heatmap is a genuine rank-3 outer-product sum next to a dense random update, rendered at a schematic 48x48 with production dimensions annotated.
Memory tools stop at retrieval. Fine-tuning products change weights with whatever gets uploaded and have nothing deciding what is true or current. Sapience is the only architecture with both halves: the store feeds context at inference and also renders the curriculum for nightly weight updates. Because the curriculum is rendered from current knowledge, adapter-trained retention landed 47 to 75 points above the no-adapter control at matched-or-better acquisition (two model families, three seeds, controlled and pre-registered; a lab result, not what ships today), structured-object replay beat the same content as raw prose by 10 to 12 points at a matched token budget (three seeds, pre-registered), and replaying the supersession-filtered store into weights leaves 0 of 81 stale facts (vs 14 of 81 unfiltered). The result is measured in the lab; what ships today is the store, retrieval, and consolidation over a frozen reader.
+32.2 points, item-paired, McNemar p=1.5e-5; five judges agree on every row. One deterministic run per block, so the robustness axis is item count, not seeds. On this task family a tuned memory layer ties us at the same budget; our +25.3pp over a tuned memory layer is on LongMemEval-S, a different benchmark (A7). A7 maps every multi-hop number in this deck to one row.

Accuracy and cost stay flat as the stored history grows from 2M to 100M tokens (81.0 / 80.0 / 80.0 / 80.0 / 79.7 at 2M / 5M / 10M / 30M / 100M; 3 seeded runs, n=300 per scale, official scorer), two orders of magnitude past any frontier window. The mechanism is not a bigger window: the reader sees about 685 tokens per question at every scale. Same figure as spnc.ai/evidence.
A different task family: our internal 10M-token corpus (populated gold, held-in cells), answering vs a frontier model with no memory. Sapience matches strong retrieval's accuracy from ~33x fewer evidence tokens per query (n=24).

Once the store exceeds the reader's window, dumping it into context collapses: 78% → 42% → 13% as it grows, all three seeds. Retrieving from a persistent store instead holds at 100%, from ~961x fewer tokens per answer at the largest store.
Retrieval over a persistent store holds at scale; dumping the store into context does not. Measured here on lookup probes, 3 seeds; the structure-specific gains are a separate study.
The multi-hop numbers in this deck are different tasks; the master table maps each to one row. The 10M corpus is internal and populated-gold; cross-vendor detail, confidence intervals, and methodology in the appendix or on request.
Each comparison runs the same reader, the same judge, and the same items across every cell. Same task in, same evaluation out. Differences come from the architecture, not from a swapped reader or an easier item set.
No single-seed claim is a headline. Headline cells run three seeds (HALO_SEED 42, 123, 456) with stochasticity on. A positive on one seed is treated as provisional and labeled as such until it survives multiple seeds. A one-seed lift that dissolves under additional seeds is not reported.
Stratified calibration holdouts are reserved for threshold and variant selection and are excluded from every headline aggregate. We never report performance on the items that set the threshold.
A second judge audits a subset of every scored run. On the cited cells, cross-judge agreement sits in the κ 0.97–1.0 range. Disagreement rows are re-adjudicated, not silently dropped.
Paired claims use item-paired tests (McNemar) on the same items scored under both conditions, not a comparison of two independent rates. The 1M matched-reader pair is item-paired at n=90, same reader both arms; the headline BABILong figures are official-scorer leaderboard cells.
Every number in this deck traces to a result file and the generator script that produced it. Nothing here is hand-entered from memory.
Since this deck first circulated: retrieval baselines at 10M landed (matched accuracy from ~33x fewer evidence tokens), the real-novels benchmark cleared multi-seed (38.6% vs 20.1%), and the MuSiQue cell now has a runnable reproduction package (available under NDA). In validation now: additional seeds on the 1M matched comparison, powered n≥30 frontier cells at full 1M windows, and a scale-fair discovery test: cross-domain surfacing over a corpus grown past any frontier window. Results go to the data room as they clear.
Adoption rule: a candidate wins only on a real margin with non-overlapping confidence intervals. Negative and null results are kept, not massaged.
| Reader | qa3 3-hop accuracy | n |
|---|---|---|
| Sapience (Sonnet 4.6 reader) | 66.7% | 90 |
| Sonnet 4.6 alone, raw 1M window | 34.4% | 90 |
Item-paired at n=90, same reader both arms, Opus primary judge with four cross-vendor judges in agreement on every row; one deterministic run per block (seed 42), so item count is the robustness axis. Other frontier readers at 1M and below: the pilot cells on the A4 figure, directional only, never quoted as powered 1M values.
| Fable 5 alone @ 660K | 64% | 50 | dual-judge, CC harness |
| Sapience + Sonnet @ 10M | 45.0% | 3 seeds | answerable, non-abstention |
| Sapience + Fable @ 10M | 41.7% | 24 | early / provisional |
| Fable 5 alone @ 10M | 0% substantive | 4 | n=4 provisional; 25% pooled incl. abstention-credit |
| Sonnet alone @ 10M | 0% substantive | – | full 1M window, gold in-context: 0 of 20 on real questions; 12.5% pooled incl. abstention-credit |
Beyond one million tokens the frontier readers cannot run; given its full 1M window with the gold in context, the same model scores 0 of 20 on real questions (12.5% with abstention credit), and 0% where the evidence lies beyond the readable window (window arithmetic). Pooled figures credit correct abstentions. Sapience is the only configuration here that still answers at ten million tokens.
Tom writes a note with his Claude: perovskite cell stability, recorded August 14, 2026, signed tom-de. His colleague asks a question in her ChatGPT and his note answers it, name attached, honest about its limits. His note, her answer, across vendors, with provenance intact. (Demonstration brains; the mechanism is live.)
Supporting row: LongMemEval-S 89.47% held-in, n=399 (canonical); in the matched whole-product pair, 88.5 vs a tuned memory layer at 63.2 (+25.3pp, p < 1e-16). We treat this as a side effect: a system that learns across time should also win memory benchmarks, and it does.
Three different multi-hop tasks: our internal 10M-token corpus (held-in cells), the 1M-token haystack (BABILong qa3), and open-domain QA (MuSiQue). At 1M the comparator is the same frontier reader over the raw window (66.7 vs 34.4, item-paired n=90); the main-slide BABILong figures (97.2 average at 10M; 80 on qa3 flat to 100M) are the official-scorer leaderboard cells. Each headline elsewhere in the deck maps to exactly one comparison here.
The one benchmark we trail is the one task memory cannot help: single-document QA, where every fact is already in the prompt and there is nothing to retrieve. We are behind exactly where our architecture is not supposed to matter.
60 of 90 with Sapience vs 31 of 90 reading the window directly: +32.2 points, item-paired, McNemar p=1.5e-5; Opus primary judge, four cross-vendor judges agree on every row. Same task, same input, same reader, same judge.
One deterministic run per block (seed 42); the robustness axis is item count, not seeds. Scope is one frontier model reading the same window, not a claim against all frontier models; the cross-vendor pilot cells on the A4 figure are directional only.
At 10M tokens: with Sapience 45.0% non-abstention (27 of 60, 3 seeds; 44.4% pooled); the same model given its full 1M window with the gold in context, 0 of 20 on real questions (12.5% pooled, the correct answers being abstentions); 0% only where the evidence lies beyond that window (window arithmetic, stated as such). The architecture is the only manipulated variable.
Non-abstention n=60 (n=72 pooled), 3 seeds, non-overlapping Wilson intervals; the bare model on its native window, 2 seeds, 10.4% pooled and 0 of 40 non-abstention; the full-1M-window run is one deterministic seed. Approved runs 2026-06-19 (close-out) and 2026-06-21 (full window). Calibration holdouts excluded from every headline aggregate (see Methodology).
Held-out accuracy from training-time consolidation (LoRA, into weights); identical inference in all arms, so the lift is not test-time compute. 95% CI +3.16 to +8.87, 10 of 10 seeds, replicated on 5 fresh seeds. Protocol in Discovery by Dreaming (arXiv 2607.16256).
Multi-hop QA over natural Wikipedia text, exact match, 3 seeds: beyond the window the Sapience ladder holds 43 to 28 percent where a truncated full-context reader falls 19 to 0. Within the window, iterative bridge retrieval adds 32 to 37 points over full context with the DeepSeek reader (the two locked cells, p<1e-4 and p=1.4e-8); a second reader family reproduces the direction at 17 to 21 points. Same ladder as spnc.ai/evidence.
What is current in an evolving codebase: currency and supersession questions over the real git history of four production repositories (Django, LiteLLM, huggingface_hub, vllm). Gold derived mechanically from git, no LLM judge. Sapience 88% with a frontier reader and 82% with an 8B open reader, vs a tuned memory layer 51%, agentic git-grep 29%, a running notes file 25%, flat RAG 18%; same model, same questions. The margin is architectural: a vector store over the same update stream holds several near-identical embeddings of the same setting with nothing marking which is current; the Sapience store knows which value is current. Same set as spnc.ai/evidence.
NoCha-style pair accuracy on public-domain classics (not the 2023+ NoCha set, which is unpublished), beyond the window, 3 seeds, 63 book pairs.
| Free | The open client and a capped free network. The adoption on-ramp. |
| Pro | $29 per person per month. |
| Team | $59 per seat, led by the builder wedge. |
| Enterprise | $50K to $500K per organization. |
| Network | A royalty when a connection we surface converts into a licensed collaboration. |
Revenue starts as subscription, seats and organization licenses; down the road a larger share may come from royalties on transacted value between researchers and industries. Each tier feeds the next: free users become the pool that Pro, Team, and Enterprise draw from.
Seven patent families, spanning structured attributed knowledge objects, matching and discovery, write-time extraction, consolidation and episodic replay, and human-AI collaboration. The continual-learning mechanisms are in filing now; priority receipts available in the data room.
Competing methods are published and open; ours are filed.
“It’s amazing to have one brain’s memory infuse all my sessions: my Claude Code and web sessions are fully aware of each other now, and they talk to my Codex sessions too, across providers. That brain is constantly compounding, and it augments my own: it proactively suggests new ideas and connections.”
“I could swap out my current engine and use yours. The cross-domain connections are things I would never have found on my own.”
“The first two papers it surfaced were spot on for my research. It is giving me ideas on what to do next.”
“The more collaboration across the world, the faster science progresses. I one billion percent see the benefit.”