Private federated learning for materials R&D

The training data for materials AI does not exist.It is locked inside companies that compete.

Every group working on batteries, catalysts and semiconductors holds experimental data that would make everyone’s models better if it were combined — and none of them will hand it over, because that data is the company. So each one trains alone on a dataset too small to matter, and the field moves at the speed of the smallest silo.

We train a shared model across those silos without the data ever leaving its owner.

Pre-seed · pre-product · one person · nothing shipped yet

The mechanism

Nothing crosses the boundary except a masked sum

Each participant trains on its own hardware, behind its own firewall. Only model updates leave, and they leave under a mask that cancels in the sum — so the aggregator sees the total and never an individual contribution.

YOUR INFRASTRUCTURE — RAW DATA NEVER CROSSES THIS LINEHolder Abattery lab · trains locally+rHolder Bcatalyst group · trains locally+rHolder Csemiconductor R&D · trains locally+rMASKED UPDATEAggregatormasks sum to zerosees the total onlyPooled modelbetter than any silo aloneRETURNED, PLUS A PRIVATE HEAD FINE-TUNED ON YOUR DATA
Federated training with secure aggregation. The masks are generated pairwise and cancel exactly in the sum, so the aggregator can compute the total without ever seeing a single participant’s update.
01

The data never moves

Not encrypted in transit — never transmitted. Raw measurements stay inside the organisation that produced them, on infrastructure it controls.

02

Updates are masked before they leave

Plain federated learning is not private: updates leak, and gradient inversion can reconstruct training examples. Secure aggregation is the difference between a promise and a guarantee.

03

Each participant gets a private model

Fine-tuned on their own data, measurably better than anything they could have trained alone. That gap is the entire product, and it is the number we are measuring first.

What you keep

You own your model. We maintain the foundation it stands on.

The ownership line sits on the technical seam that is already there. The general representation is learned from everyone and no single group would ever build it — no one lab’s task justifies its generality. Everything task-specific is different for every participant by nature, and it is yours.

Yours, outright

  • Your raw data — it never left
  • Your fine-tuned model and every output from it
  • Every discovery you make with it. We claim nothing downstream
  • A perpetual offline deployment
  • An escrow claim on the shared encoder

Held by us, as custodian

  • The shared encoder trained across everyone
  • The pipeline, and the harmonisation across labs that makes it work
  • The evaluation — the only place the pooled comparison can be computed

Somebody has to hold the shared part, and it structurally cannot be one of the participants — no company accepts that a competitor holds the asset built from its own data. That is a position, not a technology.

Why this is a company

The last time this worked, it was a research project — and it stopped when the funding did.

Continuity

A consortium is a project with an end date. A pooled model only stays valuable if it keeps training. A consortium cannot sell itself continuity.

Neutrality

A consortium of rivals still needs a trusted operator, and standing one up is a governance project none of them wants to run.

The mechanism is open

We publish the method and the code. That is deliberate: participants can verify the custodian cannot peek. What accumulates is the model, and it cannot be rebuilt from the code that built it.

Where we actually are

Most of this does not exist yet, and saying so is the point.

CompanyNot incorporated
FundingNone
CustomersNone. Nobody has signed anything
TeamOne person

What exists is the research, one peer-reviewed publication behind the technical thesis, and a benchmark under construction. The next milestone is a working two-party private training run on real materials data. If that run shows pooling does not help enough to matter, we will publish that result and stop — the threshold is written down in advance.

What is proven, and what is not

The first participants are academic groups

They have data, they want co-authorship, and they have no IP counsel to satisfy. If that is you — or if you think the whole idea is wrong and can say why — that is the conversation we want.

Start a conversation