How Artemis works
Artemis combines two ideas: reduce a problem to what is actually true, and let AI agents run many small experiments against a judge they cannot change.
Start from first principles
The lab borrows Elon Musk's habit of breaking a problem down to the facts that are known to be true, and his five-step "algorithm" for improving any process. Applied to scientific discovery, it reads like this.
Make the requirements less dumb. Every rule the lab follows is traced to the person who set it, so it can be questioned. The record is kept in a public decisions file.
Delete. The expensive step in materials discovery is the accurate calculation, not the search. So the lab removes everything that does not help it choose which calculation to run next.
Simplify, then optimize. A simple rule a scientist can read beats a black box with the same score. The agents keep a new idea only if it earns its added complexity.
Accelerate the cycle. Each experiment is scored in seconds, so the agents can try many ideas in an afternoon.
Automate last. The loop runs on its own only after the judge, the data and the rules are fixed.
Three files, after Karpathy's AutoResearch
Andrej Karpathy's autoresearch lets an AI agent improve a model by editing one file while a second, frozen file keeps score. Artemis generalizes that pattern from training language models to any measurable scientific question.
| File | Who changes it | What it holds |
|---|---|---|
harness.py | Humans only | The data check, the candidate pool, the success threshold, the metric, the time limit and the experiment log. It is fingerprinted, so any change shows. |
experiment.py | The agents | One function that ranks the candidates. This is the current best idea, and each trial is a copy with one change. |
program.md | Humans only | The instructions: setup, the experiment loop, the rules, the simplicity criterion, starting ideas, and "never stop" until the budget is spent. |
Six agents with separate jobs
The agents run together in one session of Omnigent, the open-source agent harness from Databricks. Separating the jobs means no agent grades its own work.
| Agent | Model | Job |
|---|---|---|
| Compiler | Claude Opus 5.5 | Leads the run, turns questions into experiments, keeps or discards ideas, runs the final test once |
| Scout | Claude Haiku 4.5 | Fast literature and database searches, citing only what it actually found |
| Planner | Claude Opus 5.5 | Reads the experiment log and proposes the next few ideas, each with a hypothesis |
| Experimenter | Claude Opus 5.5 | Implements one idea and scores it; several work in parallel |
| Skeptic | Claude Opus 5.5 | Reproduces every apparent gain and checks it for leakage, overfitting and needless complexity |
| Scribe | Claude Opus 5.5 | Writes the report from the log and the final test, with every source |
Guardrails
Spending caps and tool-call limits are enforced by Omnigent's own policies. The Scout, Planner and Skeptic cannot edit files, and destructive commands are blocked. The harness refuses any experiment that reads files, opens network connections or reaches around it. It allows a fixed number of experiments per run, and it lets the held-out test be used only once. Every result records the fingerprints of the harness and the data.
How the pieces connect
A question typed on this site travels to an agent sandbox and back as a live stream.