An AI lab that runs its own experiments and shows its work.
Artemis turns a scientific question into a measurable experiment. A team of AI agents runs it on public data, and everything is published: the code, the data, the failures and the final report.
In its first run, the agents built a way to choose solar-cell materials to calculate. On data they had never seen, it found every excellent candidate in 8 expensive calculations. The textbook rule needed 20, and random picking would need about 344.
The first result
The lab looked for solar-cell materials made only of earth-abundant elements, with the most toxic heavy metals excluded. A material counts as excellent if its calculated efficiency limit is at least 30 percent and it is close to stable. That calculation is expensive, so the order you run candidates in decides how quickly you find anything.
The agents tried 19 ideas, kept 4, and rejected their own highest-scoring trial as a fluke of the small practice set. Four hits on the final test is too few to trust alone, so afterwards a rebuilt copy of their method was tested on 20 larger pools. It beat the textbook rule in all 20 and the corrected-gap rule in 19.
Held-out test: expensive calculations needed to find all 4 excellent absorbers among 429 candidates. Fewer is better.
- Artemis method
- 8
- Textbook rule: stable first, then band gap nearest 1.34 eV
- 20
- Random order, expected
- 344
Expensive calculations needed to find 10 excellent absorbers
20 random pools of 2,204 earth-abundant materials. Each dot is one pool; the bar marks the median. Fewer is better.
How a run works
Each run follows the same five steps, done by agents with separate jobs, so no agent grades its own work.
Frame the question
CompilerThe question becomes a fixed test with a success threshold, a metric and a held-out set, written down before any idea is tried.
Read the field
ScoutQuick searches of the literature and materials databases, citing only sources that were actually found.
Experiment in waves
Planner and ExperimentersIdeas are proposed, then implemented and scored in parallel against a frozen judge.
Challenge every gain
SkepticAny improvement is re-run, read for leakage and tested again before it is kept.
Test once and report
Compiler and ScribeThe final method meets the held-out data exactly once, and the report records what failed as well as what worked.
Why the numbers can be trusted
A frozen judge
The scoring code and data are fingerprinted. Every result records those fingerprints, so any change would show.
Rules set in advance
The success threshold and metric were fixed before the first experiment, in code the agents cannot edit.
One look at the answer
The held-out test runs exactly once. The lab cannot try again for a better number.
Everything is public
The report, experiment log, code and data are open. A rebuild of the final method matched the lab's practice score exactly.
Have a question worth testing?
Describe it, and the lab's Compiler drafts an experiment you could run: what to measure, which public data to use, what to compare against, and whether Artemis can run it today.