An AI lab that runs its own experiments and shows its work.

Artemis turns a scientific question into a measurable experiment. A team of AI agents runs it on public data, and everything is published: the code, the data, the failures and the final report.

In its first run, the agents built a way to choose solar-cell materials to calculate. On data they had never seen, it found every excellent candidate in 8 expensive calculations. The textbook rule needed 20, and random picking would need about 344.

An illustrative rock-salt crystal lattice. Drag to turn it.

The first result

The lab looked for solar-cell materials made only of earth-abundant elements, with the most toxic heavy metals excluded. A material counts as excellent if its calculated efficiency limit is at least 30 percent and it is close to stable. That calculation is expensive, so the order you run candidates in decides how quickly you find anything.

The agents tried 19 ideas, kept 4, and rejected their own highest-scoring trial as a fluke of the small practice set. Four hits on the final test is too few to trust alone, so afterwards a rebuilt copy of their method was tested on 20 larger pools. It beat the textbook rule in all 20 and the corrected-gap rule in 19.

Held-out test: expensive calculations needed to find all 4 excellent absorbers among 429 candidates. Fewer is better.

Artemis method
8
Textbook rule: stable first, then band gap nearest 1.34 eV
20
Random order, expected
344

Expensive calculations needed to find 10 excellent absorbers

20 random pools of 2,204 earth-abundant materials. Each dot is one pool; the bar marks the median. Fewer is better.

Median calculations: Artemis method 14.5, corrected-gap rule 26, textbook rule 47.5, random order about 1,161. The textbook rule puts stable materials first, then picks band gaps nearest 1.34 eV. The corrected-gap rule, which the agents derived from the training data, aims at the right gap for this kind of calculation. This check came after the official test and used a rebuilt copy of the method; read the details.

How a run works

Each run follows the same five steps, done by agents with separate jobs, so no agent grades its own work.

  1. Frame the question

    Compiler

    The question becomes a fixed test with a success threshold, a metric and a held-out set, written down before any idea is tried.

  2. Read the field

    Scout

    Quick searches of the literature and materials databases, citing only sources that were actually found.

  3. Experiment in waves

    Planner and Experimenters

    Ideas are proposed, then implemented and scored in parallel against a frozen judge.

  4. Challenge every gain

    Skeptic

    Any improvement is re-run, read for leakage and tested again before it is kept.

  5. Test once and report

    Compiler and Scribe

    The final method meets the held-out data exactly once, and the report records what failed as well as what worked.

Why the numbers can be trusted

A frozen judge

The scoring code and data are fingerprinted. Every result records those fingerprints, so any change would show.

Rules set in advance

The success threshold and metric were fixed before the first experiment, in code the agents cannot edit.

One look at the answer

The held-out test runs exactly once. The lab cannot try again for a better number.

Everything is public

The report, experiment log, code and data are open. A rebuild of the final method matched the lab's practice score exactly.

Have a question worth testing?

Describe it, and the lab's Compiler drafts an experiment you could run: what to measure, which public data to use, what to compare against, and whether Artemis can run it today.