← Back to Data Room
CRX·2026·08  /  Report No. 01
Technical Report No. 01  /  The Sapience Program
Confidential · Prepared for Investor Review · August 2026
Seed Round  ·  August 2026

Sapience Labs


Continual learning, for science.
How fast could science move?
Discovery runs as a system of learning loops, and we ask that question of the whole system, not one tool inside it. Every frontier model is trained once and then frozen; we build the layer that keeps learning from your work, in your control, across every AI you use, closing the roughly 72 percent of discovery time that passes before the right knowledge reaches the right person, in academia and industry alike.
The intelligence that keeps learning.
Sapience Labs · built on the open CoreTx knowledge-object standard
Live · brain.coretx.ai · beta cohort forming · 7 papers, five on arXiv · seven patent families
01 / 10
02
What if?

Researchers told us why they hold back. We built to those reasons.

Two years of conversations with working scientists. The same reasons came up again and again.
they saidIts memory is a profile of my preferences, not the state of my work. Every session still starts from scratch on the things that matter.
What if it knew what you settled last Tuesday, and why: the decision, the evidence, the alternatives you already ruled out?
they saidEverything I teach it trains someone else’s model, for free, in their walled garden, with my name stripped from it.
What if what you taught it stayed yours: private by default, portable across every AI, credited wherever it traveled?
And the wish nobody thinks to say out loud
When you choose: what if two people could think together through their AIs, and what one lab learns reached the person stuck on it in days, not a decade?
What if none of this were a promise?
Live today at brain.coretx.ai. The rest of this deck is the evidence. One of these is still a promise: on most days the person working on the same problem in another field is not in the network yet. That part we can only build with you.
02 / 10
03
The Category

We build the third component of the AI stack, the one that learns while it works.

The model and the harness are the first two components; ours is the one with its own state and its own learning rules. The model stays frozen and swappable, and the system around it learns while it works: a world model of your research.
Where it sits
THE LEARNING SYSTEM what we build Learning rules what to keep, what replaces what, what consolidates overnight Attribution every claim signed and credited to its author Connection your knowledge matched to the people who need it Storage and recall what memory tools sell THE HARNESS prompts, tools THE MODEL any vendor, swappable
Memory tools supply one layer of the component. The rest is the architecture.
03 / 10
04
Already built · matched protocol, multi-seed

The evidence, with its controls.

BABILong qa3 from 4K to 100M tokens of stored history: the language model alone falls past 128K while Sapience holds a flat band near 80 out to 100M
BABILong qa3, 4K to 100M tokens. Left panel under the dual judge, right panel under the official scorer. The language model alone falls past 128K; with Sapience the line holds near 80 out to 100M. Same figure as spnc.ai/evidence.
The unit everything here runs on: a knowledge object.
A knowledge object
α-synuclein aggregates under → oxidative stress
type: finding · confidence 0.88
evidence
linked to the exact passage and dataset it was drawn from
reasoning
carries the condensed derivation behind the claim, not the transcript
edges
typed links to 4 related findings
validity
stays current as later findings replace earlier ones, and the old versions are kept
signed
to its author, verifiable (ORCID)
Selected fields, illustrative. The full property schema and the matching layer are proprietary.
97.2

#1 overall on BABILong at 10M tokens, official scorer; best published 76.6

80%

multi-hop (qa3), flat from 2M to 100M tokens; ~685 tokens read per question

88% vs 51%

what is current in an evolving codebase: Sapience vs a tuned memory layer; grep 29%

Coding, real git history: on what is current in an evolving codebase, Sapience answers 88 percent (82 on an 8B open model) where a tuned memory layer scores 51, grep 29, a notes file 25. Gold is derived mechanically from git, no judge. And within a window on a static corpus, strong retrieval ties us: we pre-registered that null ourselves.
BABILong figures come from a task-adapted retrieval pipeline, in the same sense that the fine-tuned leaderboard entries are task-trained; our zero-adaptation audit (24.7) is published beside the crown. The 97.2 average is carried by extraction-solvable tasks; qa2 and qa3 are the reasoning-bearing ones, and the per-task table is on the evidence page. Every number on this slide matches spnc.ai/evidence.
04 / 10
05
In the field

Live today, with the honest numbers.

“It’s amazing to have one brain’s memory infuse all my sessions: my Claude Code and web sessions are fully aware of each other now, and they talk to my Codex sessions too, across providers. That brain is constantly compounding.”Alex · Princeton
An earlier version of our co-scientist reached researchers in nine university labs; the Sapience beta cohort is forming now. We are customer zero: this raise was run through the product. Overnight consolidation is live, and sharing between two people is live.
The live product: a researcher's brain, mapped; a cross-domain question answered from its own cited knowledge objects
Usage, as instrumented this week · printed as measured
110

on the waitlist (unique real emails)

16

installs since Aug 19 (distinct tokens; 35 ever)

31

brains with content (at least one note)

18

active in the last 7 days (a write, a retrieval, or a new note)

12 / 4

connect codes minted / redeemed

4 / 2

share links active / joined this week

Counts from the hosted app as of Sep 3, 2026; test tokens excluded; installs before Jul 24 are unattributable. “About 100 beta users” is true only as the waitlist; the product numbers are 31 brains with content and 18 active this week. Notes per active brain in the last 7 days: median 937, range 0 to 34,256; the top end is bulk ingestion, not typing.
Small numbers are printed small. Instrumentation started the week of August 31; the first retention curve follows when there are enough weeks to draw one.
05 / 10
06
Where the time goes

Most of the delay between a problem and its solution is a connection left unmade, and the connection is cross-domain.

“…publication has been extended far beyond our present ability to make real use of the record.”Vannevar Bush, 1945
Generation is scaling. Travel is not. Across 48 of 50 documented cross-domain breakthroughs (two excluded for no measurable lag), a median 72 percent of the delay from problem to solution was reducible in principle: the winning approach already existed, in another field, before validation began, and the connection was never made. That 72 percent is the target this company is built toward.
29 / 155

realized imports findable a decade early (a floor; top-40 slate, abstract-level ranks)

12.5 yr

median time the true source sat surfaceable before import, across those 29 cases

Each dot is a real cross-domain discovery; the key source existed in another field's literature for a median of 12.5 years
Source lead time · for the 29 of 155 realized imports whose true source sat in the top-40 slate, it had been surfaceable a median 12.5 years before anyone imported it; every one of the 32 in-budget rows had more than a year of lead. Source surfacing, not forecasting.
06 / 10
07
Why Us

Other labs raise on a thesis. Ours is running.

Full-time team

Oliver Zahn PhD, Founder & CEO

Astrophysics PhD work at Harvard. Directed UC Berkeley's Center for Cosmological Physics. Head of Data Science at Google, Senior Data Scientist at SpaceX. Founder & CEO of Climax Foods.

Dr. Simran Chana

Surgeon and neuroscientist. Director, Frontier Technologies Lab, Cambridge.

Hashim Piracha

Engineering Lead. Machine learning and materials science, UC Berkeley.

Mason Rodriguez Rand

Lab Integration Lead. Molecular engineering, UChicago; founded a DOE-funded energy startup.

Senior Research Scientist

Google DeepMind. Speak with them under NDA; joins full-time on close.

Senior Research Scientist

Google DeepMind. Speak with them under NDA; joins full-time on close.

Senior Researcher

OpenAI. Speak with them under NDA; joins full-time on close.

Investors can speak with the incoming frontier-lab researchers under NDA today; they join full-time when the round closes. Six papers, five on arXiv; seven patent families, with the continual-learning mechanisms in filing now.

What sets us apart is the evidence behind the claims: matched protocols, multiple seeds, published kill conditions, and nulls reported on the slide. Diligence gets a reproduction package.

Advisors and investors

Peter Norvig

ex-Director of Research, Google; co-author, AI: A Modern Approach

David Eagleman

Neuroscientist, Stanford.

James Evans

Science of science, UChicago.

Matias Zaldarriaga

Cosmologist, Institute for Advanced Study.

Shervin Pishevar

Sherpa Capital; early Uber, Airbnb.

Ravin Kumar

Google DeepMind.

Tom Zschach

CIO, SWIFT.

Julia Yan

Former Head of Growth, TikTok.

Dean Banks

Former CEO, Tyson Foods.

07 / 10
08
Timing

The system-level learning space is open now, and this work is measured, filed, and in the field.

The two funded bets will reach for system-level learning once their own pillars mature. Today this space is ours: measured, filed, and in the field. Capital-efficient by design: we do not pre-train, and an 8B open model on the store beats a model 30x its size reading the raw context.
~12% vs ~70%

compute share of this raise vs a pretraining lab's

30x

an 8B open model on the store beats a model 30x its size on raw context

$3T → $270B

R&D inside the wedge; enterprise-AI spend at the reach

How it monetizes

The wedge is research organizations, entered from the bottom up: researchers adopt the system inside the AI sessions where their work already happens. Free client, Pro and Team seats, Enterprise licenses; revenue starts as subscription, royalties on transacted value later. Signed LOIs with two major scientific publishers sit outside this plan while those partners are paused. The full pricing table is in the appendix.

Sovereign where it needs to be

The entire system can run on customer hardware with no data leaving the building; for regulated industry and enterprise outside the United States that is often the deciding requirement, and no closed-weights lab can offer it.

08 / 10
09
The Horizon

Sapience turns what we each know into what we all know. Owned, never taken.

AI today generates. What it cannot yet do is connect. Sapience is the connective layer: one living memory that connects your sessions to each other, your AIs to each other, you to your AIs - and, with consent, you to other people.
This isn’t several products stapled together. The same knowledge object that gives one person’s AI memory is exactly the object that lets two people’s AIs connect. Build the memory and the network is latent in it. Sharing is live today: connecting two brains takes one sentence and one yes.
Ultimately, the network is the thing we are building.
Single-player from day one. Networked by design.
09 / 10
10
The Ask

Every model ships frozen. Ours keeps getting more powerful.

The ask. A seed round weighted to senior research and engineering hires, funding three milestones: the first paid cross-organization deployments, the weight path productized at the scale this deck states, and the network reaching self-sustaining density. Sapience is at its most useful on day one, alone; the network is not the price of admission but the compounding on top.

ALLOCATION · 40% engineering & product · 18% research · 22% network & go-to-market · 12% compute & infrastructure · 8% operations & reserve

The core plan is people, on roughly 24 months of runway; we serve open weights rather than pretraining, so compute is ~12% of the raise rather than ~70%. Sizing detail in the appendix.

The question this round exists to answer: how fast could science move?
You own your intelligence. Sapience makes it compound.
Oliver Zahn
oliver@coretx.ai
Sapience Labs
Investors buy Sapience Inc, which owns or exclusively licenses the architecture filings; CoreTx Inc stewards the open knowledge-object standard.
Infrastructure for Humans and AI to Discover Together
10 / 10
Appendix A1 · supporting record
Appendix A1
Appendix A1 | The Catch-22

Frontier science needs a model trained today, not last year.

A model trained on yesterday's internet cannot lead a field that changes today.
"In today's systems, we train them, and then they're kind of frozen and put out into the world. What you'd like is for those systems to continually learn online from experience."
Demis Hassabis · India AI Impact Summit · 2026
01The knowledge that decides the next breakthrough is being created right now, unpublished and never scraped. Frontier models do not contain it.
02Scientists will not pour unpublished work into a frontier lab. OpenAI launched Prism for scientists in January 2026 and closed it by April over exactly this trust wall (Decrypt).
03And retraining cannot keep up: a frozen model refreshes once or twice a year; science moves every day.

Sapience resolves both. It learns continuously with no retraining, and attribution at write time lets scientists contribute while keeping ownership and credit. One mechanism serves scientists and builds the pre-publication dataset frontier AI currently lacks.

A1 / 21
Appendix A2 · supporting record
Appendix A2
Appendix A2 | Mechanisms

A reasoning and discovery architecture, with a swappable core.

The reasoning is ours and the model plugs in. Swap in any frontier or open model and the mechanisms are the same, because they live in the system around the core. Store-and-recall is the commodity that memory tools sell. Consolidation, replay, routing, and cross-domain edges are not, and they are what produce the gains on the next slide. The attribution tier at the top of the stack is the one layer biology never needed: it gives every claim a provenance so knowledge can travel between minds with credit intact.
Nine mechanisms a memory layer does not have
Prediction-error gated encoding
writes hard on what is surprising, drops the redundant
Sleep consolidation and replay
reorganizes and links what it has seen, offline
Reconsolidation on recall
a memory becomes editable when it is used, and updates
Vividness decay and reweighting
forgets the stale, keeps what is still in use
Competitive inhibition
the strongest relevant memory wins the context
Belief revision
contradictions are versioned, never silently overwritten
Contextual reinstatement
recalls what was true in the moment a thing was learned
Provenance on every claim
source and author travel with the fact
Cross-domain matching
connects a result here to a need there, by shared mechanism
Nature solved continual learning twice · science already runs this architecture
A brain
hippocampus, fast episodes
cortex, slow consolidation
sleep, replay
forgetting, renewal
Science
journals, fast episodes
textbooks, slow canon
review and synthesis
retraction and paradigm shift
An AI that joins
a store of typed atoms
weights
overnight consolidation
supersession

The knowledge object itself carries what the mechanisms need: provenance anchors it to reality, typed links carry how it connects, and the record stays current as findings replace one another. The nine mechanisms operate on the object itself (there is no separate database underneath).

ATTRIBUTION every claim carries its author and source, so credit is auditable AT SCALE META-MEMORY confidence, contradiction detection, validity windows PROCEDURAL gates that decide what to admit, surface, consolidate SEMANTIC consolidated abstractions, schema store = any frontier or open model plugs in here EPISODIC every claim with provenance and session context INDEXING
Plate I · The Learning System · the reasoning core is swappable; the learning system around it is the architecture, and it is ours
A2 / 21
Appendix A3 · supporting record
Appendix A3
Appendix A3 | How it writes the weights

How the store writes to weights, with the base model left frozen.

A memory tool stops at retrieval. Our consolidation layer also generates the night's training curriculum from what the store currently believes, the way replay during sleep rehearses the day. A small adapter learns it and the base model is never retrained. What ships today trains nothing (frozen models reading the store); this is the lab layer, measured on open models up to 8B under LoRA.
The system view · store to a removable adapter beside a frozen base
EPISODIC STORE typed, attributed, versioned claim · evidence · provenance finding · retrieved 14x · current old value · superseded constraint · pinned · current Write gate what is worth learning tonight Replay renderer renders the night's curriculum: new material interleaved with replay of what the store currently believes SUPERSESSION FILTER stale facts are never rehearsed superseded → discarded TRAINS TRAINED NIGHTLY Adapter LoRA · a few MB FROZEN BASE MODEL tens of gigabytes · never retrained the learned statistical knowledge of pretraining, left untouched answer = W·x + B·A·x the effective weights change every night WHAT A PURE-WEIGHTS UPDATE DOES NOT HAVE REVERSIBLE delete the adapter file and the model is what it was RETRACTABLE supersession-filtered replay leaves 0 of 81 stale facts (14 of 81 unfiltered) AUDITABLE the curriculum is inspectable text with per-fact provenance Adapter-trained retention 47 to 75 points above the no-adapter control at matched-or-better acquisition (two model families, 3 seeds, pre-registered; lab result, not shipped). Replaying structured objects beats the same content as raw prose by 10 to 12 points at a matched token budget (3 seeds, pre-registered).
The matrix view · the update lives in a rank-64 channel, bounded and removable
Two heatmaps side by side. Left: a dense full fine-tuning update where every weight moves. Right: a genuine rank-3 low-rank update, where visible repeated banding shows the change is confined to a few directions.
full fine-tuning: dense, every weight moveslow-rank: a few directions
FROZEN · NEVER TRAINED W 4096 x 4096 + TWO THIN MATRICES · TRAINED OVERNIGHT B 4096 x 64 A 64 x 4096 = WHAT IT COMPUTES WITH W′ = W + B·A full-size effect, low-rank cause W holds 16.8M numbers and is never trained. B and A together hold about 0.5M numbers, roughly 3% of the layer, and are the only thing trained each night. Delete them: the base is untouched.

The visible banding on the right is the constraint made real. A low-rank update repeats a few patterns across the whole matrix instead of rewriting it freely, so it moves the model along a handful of directions and stays removable. The store decides what flows through that narrow channel each night. The heatmap is a genuine rank-3 outer-product sum next to a dense random update, rendered at a schematic 48x48 with production dimensions annotated.

Why this is not a memory harness

Memory tools stop at retrieval. Fine-tuning products change weights with whatever gets uploaded and have nothing deciding what is true or current. Sapience is the only architecture with both halves: the store feeds context at inference and also renders the curriculum for nightly weight updates. Because the curriculum is rendered from current knowledge, adapter-trained retention landed 47 to 75 points above the no-adapter control at matched-or-better acquisition (two model families, three seeds, controlled and pre-registered; a lab result, not what ships today), structured-object replay beat the same content as raw prose by 10 to 12 points at a matched token budget (three seeds, pre-registered), and replaying the supersession-filtered store into weights leaves 0 of 81 stale facts (vs 14 of 81 unfiltered). The result is measured in the lab; what ships today is the store, retrieval, and consolidation over a frozen reader.

A3 / 21
Appendix A4 · supporting record
Appendix A4
Appendix A4 | Proof at Scale

It beats strong retrieval, and holds when the corpus outgrows the window.

Against the same reader alone, on multi-hop reasoning
BABILong qa3 at 1M · same reader both arms · item-paired n=90
66.7%

with Sapience (60 of 90)

34.4%

the same model reading the raw 1M window (31 of 90)

+32.2 points, item-paired, McNemar p=1.5e-5; five judges agree on every row. One deterministic run per block, so the robustness axis is item count, not seeds. On this task family a tuned memory layer ties us at the same budget; our +25.3pp over a tuned memory layer is on LongMemEval-S, a different benchmark (A7). A7 maps every multi-hop number in this deck to one row.

Accuracy flat as the stored history scales to 100M · BABILong qa3
BABILong qa3 from 4K to 100M tokens: the language model alone falls past 128K while Sapience holds a flat band near 80 out to 100M.

Accuracy and cost stay flat as the stored history grows from 2M to 100M tokens (81.0 / 80.0 / 80.0 / 80.0 / 79.7 at 2M / 5M / 10M / 30M / 100M; 3 seeded runs, n=300 per scale, official scorer), two orders of magnitude past any frontier window. The mechanism is not a bigger window: the reader sees about 685 tokens per question at every scale. Same figure as spnc.ai/evidence.

A separate 10M task · answering vs frontier-no-memory · internal 10M corpus, held-in cells, matched protocol
0%

frontier at 10M (field exceeds its window)

45.0%

with Sapience at 10M (27 of 60 non-abstention, 3 seeds), vs 0 of 20 for the same model given its full 1M window (12.5% pooled, the correct answers being abstentions)

A different task family: our internal 10M-token corpus (populated gold, held-in cells), answering vs a frontier model with no memory. Sapience matches strong retrieval's accuracy from ~33x fewer evidence tokens per query (n=24).

As the store grows past the window · accuracy and cost
As the accumulated store grows past the context window, dumping it into context collapses in accuracy and explodes in tokens per answer, while retrieval over a persistent store holds flat on both.

Once the store exceeds the reader's window, dumping it into context collapses: 78% → 42% → 13% as it grows, all three seeds. Retrieving from a persistent store instead holds at 100%, from ~961x fewer tokens per answer at the largest store.

Retrieval over a persistent store holds at scale; dumping the store into context does not. Measured here on lookup probes, 3 seeds; the structure-specific gains are a separate study.

The multi-hop numbers in this deck are different tasks; the master table maps each to one row. The 10M corpus is internal and populated-gold; cross-vendor detail, confidence intervals, and methodology in the appendix or on request.

A4 / 21
Appendix A5 · supporting record
Appendix A5
Appendix A5 | Methodology

Matched protocol, multi-seed, leakage-clean.

Every headline number runs on a matched protocol: same reader, same judge, same seed set, multi-seed, blind evaluation. The full method, end to end.

Matched protocol

Each comparison runs the same reader, the same judge, and the same items across every cell. Same task in, same evaluation out. Differences come from the architecture, not from a swapped reader or an easier item set.

Multi-seed validation

No single-seed claim is a headline. Headline cells run three seeds (HALO_SEED 42, 123, 456) with stochasticity on. A positive on one seed is treated as provisional and labeled as such until it survives multiple seeds. A one-seed lift that dissolves under additional seeds is not reported.

Calibration holdouts excluded

Stratified calibration holdouts are reserved for threshold and variant selection and are excluded from every headline aggregate. We never report performance on the items that set the threshold.

Cross-judge audits

A second judge audits a subset of every scored run. On the cited cells, cross-judge agreement sits in the κ 0.97–1.0 range. Disagreement rows are re-adjudicated, not silently dropped.

Item-paired statistics

Paired claims use item-paired tests (McNemar) on the same items scored under both conditions, not a comparison of two independent rates. The 1M matched-reader pair is item-paired at n=90, same reader both arms; the headline BABILong figures are official-scorer leaderboard cells.

Every number is traceable

Every number in this deck traces to a result file and the generator script that produced it. Nothing here is hand-entered from memory.

Living record

Since this deck first circulated: retrieval baselines at 10M landed (matched accuracy from ~33x fewer evidence tokens), the real-novels benchmark cleared multi-seed (38.6% vs 20.1%), and the MuSiQue cell now has a runnable reproduction package (available under NDA). In validation now: additional seeds on the 1M matched comparison, powered n≥30 frontier cells at full 1M windows, and a scale-fair discovery test: cross-domain surfacing over a corpus grown past any frontier window. Results go to the data room as they clear.

Adoption rule: a candidate wins only on a real margin with non-overlapping confidence intervals. Negative and null results are kept, not massaged.

A5 / 21
Appendix A6 · supporting record
Appendix A6
Appendix A6 | Cross-Vendor

Cross-vendor at one million tokens.

Frontier readers alone on BABILong qa3 (3-hop), reading the raw window directly. Sapience holds where every frontier reader collapses. Provisional cells are labeled.
At 1M tokens · BABILong qa3 3-hop
Reader qa3 3-hop accuracy n
Sapience (Sonnet 4.6 reader) 66.7% 90
Sonnet 4.6 alone, raw 1M window 34.4% 90

Item-paired at n=90, same reader both arms, Opus primary judge with four cross-vendor judges in agreement on every row; one deterministic run per block (seed 42), so item count is the robustness axis. Other frontier readers at 1M and below: the pilot cells on the A4 figure, directional only, never quoted as powered 1M values.

At 660K and beyond · the context ceiling
Fable 5 alone @ 660K 64% 50 dual-judge, CC harness
At ten million tokens · internal 10M corpus, held-in cells, matched protocol
Sapience + Sonnet @ 10M 45.0% 3 seeds answerable, non-abstention
Sapience + Fable @ 10M 41.7% 24 early / provisional
Fable 5 alone @ 10M 0% substantive 4 n=4 provisional; 25% pooled incl. abstention-credit
Sonnet alone @ 10M 0% substantive full 1M window, gold in-context: 0 of 20 on real questions; 12.5% pooled incl. abstention-credit

Beyond one million tokens the frontier readers cannot run; given its full 1M window with the gold in context, the same model scores 0 of 20 on real questions (12.5% with abstention credit), and 0% where the evidence lies beyond the readable window (window arithmetic). Pooled figures credit correct abstentions. Sapience is the only configuration here that still answers at ten million tokens.

Cross-vendor in the field · two people, two AIs

Tom writes a note with his Claude: perovskite cell stability, recorded August 14, 2026, signed tom-de. His colleague asks a question in her ChatGPT and his note answers it, name attached, honest about its limits. His note, her answer, across vendors, with provenance intact. (Demonstration brains; the mechanism is live.)

A6 / 21
Appendix A7 · supporting record
Appendix A7
Appendix A7 | The Numbers

The complete record.

Every headline in this deck on one figure, with the matched-protocol notes. Each benchmark measures a different thing; the labels say which.
Eight benchmarks · matched protocol · selected rows, with intervals, are on spnc.ai/evidence

Supporting row: LongMemEval-S 89.47% held-in, n=399 (canonical); in the matched whole-product pair, 88.5 vs a tuned memory layer at 63.2 (+25.3pp, p < 1e-16). We treat this as a side effect: a system that learns across time should also win memory benchmarks, and it does.

Three different multi-hop tasks: our internal 10M-token corpus (held-in cells), the 1M-token haystack (BABILong qa3), and open-domain QA (MuSiQue). At 1M the comparator is the same frontier reader over the raw window (66.7 vs 34.4, item-paired n=90); the main-slide BABILong figures (97.2 average at 10M; 80 on qa3 flat to 100M) are the official-scorer leaderboard cells. Each headline elsewhere in the deck maps to exactly one comparison here.

The one benchmark we trail is the one task memory cannot help: single-document QA, where every fact is already in the prompt and there is nothing to retrieve. We are behind exactly where our architecture is not supposed to matter.

left in on purpose.
Matched-reader pair · BABILong qa3 @ 1M · not a headline
66.7 vs 34.4%

Sapience vs the same Sonnet 4.6 reader over the raw 1M window

60 of 90 with Sapience vs 31 of 90 reading the window directly: +32.2 points, item-paired, McNemar p=1.5e-5; Opus primary judge, four cross-vendor judges agree on every row. Same task, same input, same reader, same judge.

One deterministic run per block (seed 42); the robustness axis is item count, not seeds. Scope is one frontier model reading the same window, not a claim against all frontier models; the cross-vendor pilot cells on the A4 figure are directional only.

10M-token reasoning · self-built, leakage-clean
45%

Internal 10M corpus · matched pair · held-in cells

At 10M tokens: with Sapience 45.0% non-abstention (27 of 60, 3 seeds; 44.4% pooled); the same model given its full 1M window with the gold in context, 0 of 20 on real questions (12.5% pooled, the correct answers being abstentions); 0% only where the evidence lies beyond that window (window arithmetic, stated as such). The architecture is the only manipulated variable.

Non-abstention n=60 (n=72 pooled), 3 seeds, non-overlapping Wilson intervals; the bare model on its native window, 2 seeds, 10.4% pooled and 0 of 40 non-abstention; the full-1M-window run is one deterministic seed. Approved runs 2026-06-19 (close-out) and 2026-06-21 (full window). Calibration holdouts excluded from every headline aggregate (see Methodology).

Non-synthetic, real-data benchmarks
Consolidation uplift · replicated
+6.01pp

Consolidation into weights

Held-out accuracy from training-time consolidation (LoRA, into weights); identical inference in all arms, so the lift is not test-time compute. 95% CI +3.16 to +8.87, 10 of 10 seeds, replicated on 5 fresh seeds. Protocol in Discovery by Dreaming (arXiv 2607.16256).

Multi-hop QA · real
43 to 28 vs 19 to 0%

MuSiQue, beyond the window

Multi-hop QA over natural Wikipedia text, exact match, 3 seeds: beyond the window the Sapience ladder holds 43 to 28 percent where a truncated full-context reader falls 19 to 0. Within the window, iterative bridge retrieval adds 32 to 37 points over full context with the DeepSeek reader (the two locked cells, p<1e-4 and p=1.4e-8); a second reader family reproduces the direction at 17 to 21 points. Same ladder as spnc.ai/evidence.

Coding · real git history
88 vs 51 / 29 / 25 / 18%

Repo-Evolution

What is current in an evolving codebase: currency and supersession questions over the real git history of four production repositories (Django, LiteLLM, huggingface_hub, vllm). Gold derived mechanically from git, no LLM judge. Sapience 88% with a frontier reader and 82% with an 8B open reader, vs a tuned memory layer 51%, agentic git-grep 29%, a running notes file 25%, flat RAG 18%; same model, same questions. The margin is architectural: a vector store over the same update stream holds several near-identical embeddings of the same setting with nothing marking which is current; the Sapience store knows which value is current. Same set as spnc.ai/evidence.

Published novels · real
38.6 vs 20.1%

NoCha-style

NoCha-style pair accuracy on public-domain classics (not the 2023+ NoCha set, which is unpublished), beyond the window, 3 seeds, 63 book pairs.

How it is priced · four layers, one product
FreeThe open client and a capped free network. The adoption on-ramp.
Pro$29 per person per month.
Team$59 per seat, led by the builder wedge.
Enterprise$50K to $500K per organization.
NetworkA royalty when a connection we surface converts into a licensed collaboration.

Revenue starts as subscription, seats and organization licenses; down the road a larger share may come from royalties on transacted value between researchers and industries. Each tier feeds the next: free users become the pool that Pro, Team, and Enterprise draw from.

A7 / 21
Appendix A8 · supporting record
Appendix A8
Appendix A8 | Papers & IP

A real research program, and the IP under it.

On the public record · seven public, one in review
  • Sapience: A Hybrid Architecture for Long-Context Reasoning and Continual Learning · Zenodo
  • Facts as First-Class Objects · arXiv
  • Selective Memory for AI (salience gating) · arXiv
  • Attention Is Not Retention · arXiv
  • Attribution-Native Machine Learning · arXiv
  • Constraint Gain: valuing negative results · ResearchGate
  • Discovery Stack · Nature Communications, in review
  • Discovery by Dreaming (cross-domain recombination) · arXiv
The IP position

Seven patent families, spanning structured attributed knowledge objects, matching and discovery, write-time extraction, consolidation and episodic replay, and human-AI collaboration. The continual-learning mechanisms are in filing now; priority receipts available in the data room.

Competing methods are published and open; ours are filed.

A8 / 21
Appendix A9 · supporting record
Appendix A9
Appendix A9 | Moat

What closed labs structurally cannot ship.

Not a memory tool bolted onto a model. An integrated architecture that compounds.
“Why won’t a frontier lab crush this?” The question comes up in every serious meeting, and it has a direct answer. They are racing each other on generators, and the neutral learning layer across all of them cannot be owned by one of them. The free memory inside Claude, ChatGPT, Cursor, and Glean is the near-term version of this question, and the answer is structural: their memory features are single-vendor by construction, and deeper than that: training on users is the model business, and attribution would price it. And capital does not close the gap: compute cannot buy the attributed knowledge that accumulates here.
A model built to absorb your experts into its weights cannot also compound their expertise for them. A frontier lab’s training pipelines require ingestion; ours do not. Their business puts your data on their servers; ours keeps it yours. They cannot follow without unwinding their own model.
The mechanisms · hard to copy Storing and recalling your own data is table stakes now, and better retrieval does not lead to this architecture: write order, supersession, and provenance are discarded at embedding time, and no amount of retrieval quality recovers what ingestion threw away. Memory and retrieval tools keep improving how they choose what to return; our advantage is in how we keep the knowledge in the first place. What compounds is connecting knowledge across organizations, shared as the mechanism rather than the raw result, with permission, credit to the contributor, and compensation. Attribution is also the training bet: provenance lets the system weight what it learns by how useful each claim has proven.
The structure · closed to a single-vendor lab The reasoning core is swappable, so no vendor owns us. A single-vendor lab structurally cannot offer a you-own-it substrate that spans every AI and every organization, including the ones it competes with. The mechanisms above are engineering; this constraint is not: matching it would mean unwinding their own business.
Open standard, sealed engine. The CoreTx standard spreads and becomes the way knowledge moves between AIs; the matching and learning engine stays Sapience’s.
1 Ownership + attribution signed to its author, yours 2 Share unpublished work because they own it 3 Most-current knowledge learns as science happens 4 Most wanted to join to contribute and learn from 5 More contribution smarter still, back to start The moat freshest collective scientific intelligence
Ownership is why scientists share; what they share keeps the system current.
Measured: when facts change, 98% current-answer accuracy on 108 update probes, 70 points over similarity retrieval (3 seeds); an order-aware retrieval steelman draws level, so the claim is scoped to relevance retrieval.
The moat is the collective knowledge the network accumulates; the underlying model is replaceable.
A9 / 21
Appendix A10 · supporting record
Appendix A10
Appendix A10 | Notes to the main slides

The text behind each main slide, kept off the slides.

Each main slide now carries one short body. What it used to say is here, keyed by slide, unchanged in substance.
02 What if: the other four reasons researchers gave
They said: It makes things up, on exactly the technical details that matter. What if every answer were grounded in your own verified record, and cited it, claim by claim?
They said: It reasons impressively about everything, and shallowly about the field I actually work in. What if it kept its frontier-wide reach, and knew your lab’s work down to which finding superseded which?
They said: It only ever answers. It never brings anything of its own. What if it kept thinking overnight, and came back with ideas of its own?
They said: Real problems do not fit in a context window, and managing it is unpaid work. What if context weren’t limited, and managing it weren’t your job?
03 The Category: the two funded bets, and where memory tools sit
AI has two funded learning bets: retrain the weights (personalization, one model gets smarter) and generate new data (robot labs). Both are real. Hypotheses are not cures; the world still has to answer. And neither bet can carry what a lab knows across models, sessions, and people.
Memory tools sell storage and recall for a single model. Dozens of funded teams compete there, inside one layer of the component. That crowd is our supply chain.
Our real competition is the in-weights bet. We win where the mechanism predicts we should, on currency, scale, and accumulation. We ran that comparison as one pre-registered experiment and published the kill condition.
Sapience covers both axes at once: breadth and nowness, where a frozen model gives you one at the cost of the other. Across fields, it bridges work no single context can hold and knows your lab down to which finding superseded which. Across time, it holds what you settled this morning and what a field learned a decade ago, and reasons over both.
04 Proof: the other two results, and the unit of knowledge
It knows things nobody told it. On aggregation questions whose answers exist in no single input, the write-time aggregation gate adds 22.8 points (95% CI +8.1 to +38.2) over a compute-matched counter emitting the identical sentence; exactly neutral on single-episode controls (paper v4.3). A storage system returns what you put in. This returns answers that were never stored.
Current knowledge: when facts change, 98 percent current-answer accuracy on 108 update probes, 70 points over similarity retrieval (3 seeds). Give the strongest retrieval baseline write order and it draws level, which tells us the gain comes from tracking order, not from a weak comparison.
Knowledge travels in two containers today: papers, fifteen pages, years late, with provenance no machine can read; and weights, where everything is absorbed and nothing travels out. The knowledge object is the container we built instead: one claim carrying its evidence, its reasoning, its typed links, and its author. The reasoning is worked out once at capture and reused after that, where a session log makes every later query re-derive it. That is why the store stays small and fast to reason over even at field scale, and the typed edges are what the discovery engine reasons over. The structure itself does measurable work all the way into the weights: replaying structured objects beats replaying the same content as raw prose by 10 to 12 points at a matched token budget (3 seeds, pre-registered).
05 In the field: what the product does overnight
Overnight consolidation is live: it comes back in the morning with grounded connections between your own open problems, each one citing the objects it came from.
The researcher quotes in the appendix are unedited. They describe one brain compounding across providers, surfacing new ideas and connections, and cross-domain links a researcher would not have found alone.
06 Where the time goes: the mechanism behind the 72 percent
The oldest loop in science, ideas crossing fields, still runs on decades: Boolean algebra waited 83 years for circuit design. In research first, across academia and industry, where the price of frozen knowledge is highest, and then in knowledge work generally.
Structured, cross-domain matching covers that gap. On imports that actually happened, the true source sits in a top-40 slate a median of a decade before anyone reaches for it (a floor, not a cherry-pick).
Reasoning outside the box is where current AI is weakest: it completes patterns close to what it has already seen. Structured knowledge lets the engine trade off deliberately between nearby associations and far analogies, matching by mechanism across fields. The imports in the figure are that tradeoff at work.
Science is where this matters most, because nowhere else does knowledge move this fast and travel this badly. And the discovery architecture compounds: every brain that joins adds structured, typed knowledge the matching engine reasons over, so it gets better the more it connects. One case from the figure: given a 2025 zinc-battery problem blind, the engine surfaces a 1986 solvation-chemistry source in its slate, 39 years after it was published.
08 Timing: data supply, cost structure, and what is open
The industry’s data supply curve is running out: frontier labs now bid millions for defunct companies’ internal chat logs. The scarcest training signal is expert reasoning, and the store captures it at write time: structured, current, owned by the people who produced it.
Capital-efficient by design: we do not pre-train. The cost advantage is structural. An 8B open model on the store beats a model 30x its size reading the truncated raw context. A stronger reader still helps within its own window, and either way there is no per-token rent, so no vendor owns us.
The standard is open. The network is open to join. The intelligence that runs it is ours. Revenue starts as subscription, and a larger share may later come from royalties on transacted value. Pricing is live; the table is in the appendix, since it is not what this round is priced on.
Seed round, weighted to senior research and engineering hires.
09 The Horizon: the substrate argument, drawn out
The same knowledge object that gives one person’s AI memory - a claim with its evidence, its reasoning, and a name attached - is exactly the object that lets two people’s AIs connect. Matching, consent, and attribution are properties of the substrate, not features bolted beside it.
Publishing in the age of AI: knowledge broken into small credited claims, with a feed. Contributions surface by what they solve, and anyone can plug their own knowledge into a live discovery process.
Each field joins the shared haystack by consent, with credit. A blind match keeps sharing safe: both sides learn a connection exists; the insight stays hidden until both opt in. Sharing between two brains is topic-scoped, read-only, revocable, across vendors. A verified connection between two labs is not yet on record; it is a named milestone of this raise.
Consider two materials labs that choose to collaborate through Sapience. They share one continuously learning system: it learns from their combined work as it happens, holds both labs’ full context where no window could, and surfaces connections between them, credited to whoever established each claim. And a measured path carries what it learns into the weights: in controlled, pre-registered runs, adapter-trained retention lands 47 to 75 points above the no-adapter control while learning the new domain as well or better, across two model families, three seeds each; a lab result, not what ships today (the mechanism, drawn out: Appendix A3). No frozen model offers this, at any size.
The only superintelligence that matters is personalized superintelligence: yours, grounded in a living model of your work, compounding with every mind that chooses to connect.
10 The Ask: wedge, sizing, and cost structure
Science is the wedge, not the ceiling. The substrate runs on any working knowledge, and the named second beachhead is builders: a repository’s decision history is working knowledge at its most concrete, and the coding result earlier in this deck is that market’s evidence. The wedge is the world’s researchers inside roughly $3T of global R&D a year; the reach is a billion knowledge workers, under $115B of enterprise-AI spend in 2026 heading toward $270B. Science is the hardest version of knowledge work, which is why the architecture carries to the rest without being rebuilt.
How the number is sized. The core plan is people: roughly 24 senior research and engineering hires on 24 months of runway, serving open weights rather than pretraining. Beyond the team, the round buys the two milestones that cannot wait: the weight path productized, and the network reaching self-sustaining density while the system-level learning space is still open. We hire from the same pool as the in-weights labs, at the same level, without the pretraining bill. The result is a frontier-lab raise with the compute line removed.
Capital-efficient by design: we do not pre-train a frontier model (~6N FLOPs per token, trillions of tokens, many runs). We serve open weights (~2N per token) and build the learning system on top, so compute is ~12% of the raise rather than ~70%, and the largest line is people: growing the full-time team toward ~24 across engineering, research, and go-to-market, on roughly 24 months of runway.
A10 / 21
Appendix A11 · supporting record
Appendix A11
Appendix A11 | Researcher voices

What researchers said, unedited.

The first quote is a current Sapience user. The other three used an earlier version of the co-scientist in 2026; institution labels only.

“It’s amazing to have one brain’s memory infuse all my sessions: my Claude Code and web sessions are fully aware of each other now, and they talk to my Codex sessions too, across providers. That brain is constantly compounding, and it augments my own: it proactively suggests new ideas and connections.”

Alex · Princeton University · Sapience

“I could swap out my current engine and use yours. The cross-domain connections are things I would never have found on my own.”

Researcher · University of Cambridge · earlier version

“The first two papers it surfaced were spot on for my research. It is giving me ideas on what to do next.”

Researcher · CERN · earlier version

“The more collaboration across the world, the faster science progresses. I one billion percent see the benefit.”

Researcher · Queen Mary University of London · earlier version
A11 / 21