Participate
What it actually involves
The first participants are academic materials groups — they have data, they want co-authorship, and they have no IP counsel to satisfy. Industrial R&D groups come after that, and the terms below are written with their lawyers in mind.
Nobody has signed anything yet. There is no pool to join today. What is available right now is a conversation about whether this would work for your data, and an early look at the benchmark when it lands.
What you contribute
- Data
- Made available to a training process that runs inside your own environment. It is not transferred, copied or stored by us.
- Compute
- Enough to run local training rounds. Modest — this is not a foundation-model-scale workload.
- A technical contact
- One person who can install the client and answer questions about the data.
What you get
- Your own model
- The shared encoder plus a head fine-tuned privately on your data, retrained on a regular cadence.
- The measurement that matters
- A benchmark of your pooled model against a model trained on your data alone. If that number is unimpressive, you should know it and so should we.
- Everything downstream
- Every prediction, candidate and discovery you make with it. We claim no interest in your results.
- Publication rights
- Nothing here restricts you publishing your own research. For academic groups, co-authorship on the methods work is on the table.
- A seat on the roadmap
- Which properties, which architectures, which datasets, which release cadence.
The questions lawyers ask first
Can our data be reconstructed from what leaves the building?
Updates are masked before they leave and the masks cancel in the sum, so the aggregator sees only the total. This is a computational guarantee with stated assumptions, not magic — the assumptions are written down in the agreement, in plain language, including what they do not cover.
Who owns the model?
You own your fine-tuned model and everything you produce with it. The shared encoder — the part trained across all participants — is held by us as neutral custodian, because it structurally cannot be held by one of the participants.
What if you go out of business, or get acquired by our competitor?
The encoder sits in escrow. It releases to you on our insolvency, on an acquisition by a competitor of yours, or if we fail to deliver a retrained model. We offer this in the first draft rather than waiting to be asked.
What happens to our contribution if we leave?
It stays in the model. Model unlearning is not offered and we do not represent it as feasible. You keep a perpetual licence to the last model delivered to you, and you stop contributing. We would rather say this in the first meeting than have you find it in redlines.
We compete with the other participants. Is this even legal?
Pre-competitive research collaboration is lawful and common, and this structure is unusually defensible: participants learn nothing about each other's inputs, so there is no information exchange to scrutinise. The arrangement stays upstream of products — property prediction, never formulation — and never touches pricing, capacity, output or launch timing. Specialist counsel reviews it before any multi-party agreement.
What it costs
For the first participants: nothing. Early groups are paid in priority, influence and co-authorship rather than charged, because they are the ones taking a risk on something unproven. Pricing for later members is a question we will answer when we have earned the right to.
Or tell us why this is wrong
A well-argued reason this will not work for your data is worth more to us right now than a polite expression of interest. We will say so if you change our mind.
Get in touch