About Artemis
Artemis is an experiment in making scientific discovery faster without making it less honest.
Why it exists
Much of science is a search: thousands of candidates, a measurement too expensive to run on all of them, and a few that matter. AI agents can search tirelessly, but an agent that grades its own homework will fool itself. Artemis gives the agents a fixed judge, rules written in advance, and a single look at held-out data. It also publishes everything, so anyone can check the result.
Who built it
Artemis was built by Peter Alesso with Claude for Challenge 03 of the 7th Global AI Hackathon. The lab runs on Omnigent, the open-source agent harness from Databricks. The servers and sandboxes run on Modal, this website on Vercel, and saved results in a Neon Postgres database. The agents use Anthropic's Claude Opus 5.5 and Claude Haiku 4.5.
Data and credit
The materials data come from NIST JARVIS-DFT and the JARVIS-Leaderboard, created by Kamal Choudhary and colleagues at the National Institute of Standards and Technology. Candidate cross-checks use the Materials Project, and literature searches use OpenAlex and arXiv. Crustal abundances come from the CRC Handbook of Chemistry and Physics.
The code is open on GitHub.
Questions
What is Artemis?
Artemis is an autonomous AI lab. It turns a scientific question into a measurable experiment, runs the experiment with a team of six AI agents on public data, tests the result once on held-out data, and publishes the report, code and data.
Did Artemis discover a new solar-cell material?
Not yet. Its first run produced a faster way to choose which materials to calculate, and a ranked list of earth-abundant crystals that have no SLME value in JARVIS. Each candidate still needs the expensive calculation, and then laboratory work, before anyone can call it a discovery.
How much faster is the Artemis method?
On the official held-out test (429 materials, 4 excellent absorbers, run once), the method found all four in 8 expensive calculations; the textbook rule (stable materials first, band gap nearest 1.34 eV) needed 20 and random search about 344. In a later check of a rebuilt copy of the method on 20 random pools of 2,204 materials, it needed a median of 14.5 calculations to find 10 excellent absorbers, against 26 for a corrected band-gap rule, 47.5 for the textbook rule and about 1,161 for random search, and it beat the textbook rule in all 20 pools.
Who built Artemis?
Peter Alesso built Artemis for Challenge 03 of the 7th Global AI Hackathon, working with Claude, Anthropic's AI assistant, as a coding partner. The code is open source under the MIT license.
Which AI models does Artemis use?
Claude Opus 5.5 plans, designs and runs experiments, reviews them as the Skeptic, and writes the report. Claude Haiku 4.5 runs quick literature and database searches. The agents run in Omnigent, the open-source agent harness from Databricks, on Modal.
Where does the data come from?
From NIST JARVIS-DFT and the JARVIS-Leaderboard, with crustal abundances from the CRC Handbook of Chemistry and Physics. The Materials Project and OpenAlex are used for cross-checks and literature searches.
Can I ask Artemis my own question?
Yes. On the Lab page, describe your question. The lab's Compiler drafts a testable experiment: what to measure, which public data to use, baselines, risks, and whether Artemis can run it today. Starting a run needs an access code because it uses paid computing; every finished run is published on the Results page.
How does Artemis keep its results honest?
The scoring code and data are frozen and fingerprinted. Success thresholds are fixed before any experiment. A separate Skeptic agent re-runs every apparent gain, and the held-out test is used exactly once. The final model of the first run was later rebuilt from its report; it matched the lab's practice score exactly and 14 of its 15 recommendations.
Is the code open source?
Yes. The harness, agents, website and results are on GitHub at github.com/alessoh/artemis under the MIT license.