How Artemis works

Artemis combines two ideas: reduce a problem to what is actually true, and let AI agents run many small experiments against a judge they cannot change.

Start from first principles

The lab borrows Elon Musk's habit of breaking a problem down to the facts that are known to be true, and his five-step "algorithm" for improving any process. Applied to scientific discovery, it reads like this.

Make the requirements less dumb. Every rule the lab follows is traced to the person who set it, so it can be questioned. The record is kept in a public decisions file.

Delete. The expensive step in materials discovery is the accurate calculation, not the search. So the lab removes everything that does not help it choose which calculation to run next.

Simplify, then optimize. A simple rule a scientist can read beats a black box with the same score. The agents keep a new idea only if it earns its added complexity.

Accelerate the cycle. Each experiment is scored in seconds, so the agents can try many ideas in an afternoon.

Automate last. The loop runs on its own only after the judge, the data and the rules are fixed.

Three files, after Karpathy's AutoResearch

Andrej Karpathy's autoresearch lets an AI agent improve a model by editing one file while a second, frozen file keeps score. Artemis generalizes that pattern from training language models to any measurable scientific question.

FileWho changes itWhat it holds
harness.pyHumans onlyThe data check, the candidate pool, the success threshold, the metric, the time limit and the experiment log. It is fingerprinted, so any change shows.
experiment.pyThe agentsOne function that ranks the candidates. This is the current best idea, and each trial is a copy with one change.
program.mdHumans onlyThe instructions: setup, the experiment loop, the rules, the simplicity criterion, starting ideas, and "never stop" until the budget is spent.

Six agents with separate jobs

The agents run together in one session of Omnigent, the open-source agent harness from Databricks. Separating the jobs means no agent grades its own work.

AgentModelJob
CompilerClaude Opus 5.5Leads the run, turns questions into experiments, keeps or discards ideas, runs the final test once
ScoutClaude Haiku 4.5Fast literature and database searches, citing only what it actually found
PlannerClaude Opus 5.5Reads the experiment log and proposes the next few ideas, each with a hypothesis
ExperimenterClaude Opus 5.5Implements one idea and scores it; several work in parallel
SkepticClaude Opus 5.5Reproduces every apparent gain and checks it for leakage, overfitting and needless complexity
ScribeClaude Opus 5.5Writes the report from the log and the final test, with every source

Guardrails

Spending caps and tool-call limits are enforced by Omnigent's own policies. The Scout, Planner and Skeptic cannot edit files, and destructive commands are blocked. The harness refuses any experiment that reads files, opens network connections or reaches around it. It allows a fixed number of experiments per run, and it lets the held-out test be used only once. Every result records the fingerprints of the harness and the data.

How the pieces connect

A question typed on this site travels to an agent sandbox and back as a live stream.

Artemis architecture The browser talks to this website on Vercel. The website talks to the Omnigent server on Modal, which starts an agent sandbox on Modal. The agents call Claude models and read public data from NIST JARVIS and the Materials Project. Events stream back to the browser, and finished reports are saved in a Neon database. Your browserquestion, live view This websiteVercel, Python relay Omnigent serverModal Agent sandboxsix agents, harness,fresh copy of the code Claude Public dataJARVIS, MP Saved results Your browserquestion, live view This websiteVercel, Python relay Omnigent serverModal Agent sandboxsix agents, harness, code Claude Public dataJARVIS, MP Saved results
Secrets stay on the servers: the Claude key is injected only into the sandbox, and starting a run requires an access code.